minjaechoi/qwen3p6-35b-a3b-2p02bit-r33
The minjaechoi/qwen3p6-35b-a3b-2p02bit-r33 model is an internal research checkpoint based on the Qwen/Qwen3.6-35B-A3B architecture, featuring 35.1 billion parameters and a 32768-token context length. This model utilizes a unique quantization scheme where routed experts average 2.022 bits, while other weights remain in BF16 format. It is designed for efficient inference with dequantized weights that are compatible with standard `transformers` and vLLM libraries, making it suitable for memory-constrained applications requiring a large language model.
Loading preview...
Model Overview
This model, minjaechoi/qwen3p6-35b-a3b-2p02bit-r33, is an internal research checkpoint derived from the Qwen/Qwen3.6-35B-A3B base model. It features 35.1 billion parameters and supports a substantial context length of 32768 tokens.
Key Differentiators
- Quantized Architecture: Employs a novel quantization strategy where "routed experts" are compressed to an average of 2.022 bits. All other weights are maintained in BF16 precision.
- Efficient Loading: Weights are stored dequantized in BF16 tensors, ensuring compatibility and ease of use with standard libraries like
transformersand vLLM without requiring specialized loading mechanisms. - Research Focus: Identified as an internal research checkpoint (ID: r33), indicating its role in exploring advanced quantization and model efficiency techniques.
Use Cases
This model is particularly well-suited for scenarios where:
- Memory Efficiency is Critical: The mixed-precision quantization allows for a significant reduction in memory footprint compared to a full BF16 or FP16 35B model.
- Large Context Processing: Its 32768-token context window makes it capable of handling extensive inputs for tasks like document analysis, long-form content generation, or complex reasoning.
- Leveraging Qwen Capabilities: Inherits the foundational capabilities of the Qwen3.6-35B-A3B base model, making it suitable for a wide range of general-purpose language tasks while benefiting from quantization for improved inference efficiency.