minjaechoi/qwen36-35b-a3b-2p00bit-r9
The minjaechoi/qwen36-35b-a3b-2p00bit-r9 model is an internal research checkpoint based on the Qwen3.6-35B-A3B architecture. This 35.1 billion parameter model utilizes routed experts averaging 2.00 bits, with all other weights stored in BF16 tensors. It is designed for efficient loading with standard transformers and vLLM libraries, focusing on optimized performance through quantization techniques.
Loading preview...
Model Overview
The minjaechoi/qwen36-35b-a3b-2p00bit-r9 is an internal research checkpoint derived from the Qwen/Qwen3.6-35B-A3B base model. This variant incorporates a specialized quantization strategy, utilizing routed experts that average 2.00 bits for efficiency. All other weights within the model are maintained in BF16 format.
Key Characteristics
- Base Model: Qwen/Qwen3.6-35B-A3B
- Parameter Count: Approximately 35.1 billion parameters.
- Quantization: Employs routed experts with an average of 2.00 bits, while non-expert weights are BF16.
- Loading Compatibility: Designed to load seamlessly with standard
transformersandvLLMlibraries, as weights are stored dequantized in BF16 tensors. - Context Length: Supports a context length of 32768 tokens.
Use Cases
This model is particularly suitable for:
- Research and Development: Ideal for exploring the performance and efficiency benefits of advanced quantization techniques like routed experts.
- Resource-Constrained Deployment: Its optimized bit-depth for experts suggests potential for reduced memory footprint and faster inference compared to full BF16 models of similar scale.
- Experimentation with Qwen Architectures: Provides a quantized variant for those working with the Qwen3.6-35B-A3B family.