minjaechoi/qwen3p6-35b-a3b-2p06bit-r59
The minjaechoi/qwen3p6-35b-a3b-2p06bit-r59 model is an internal research checkpoint based on the Qwen3.6-35B-A3B architecture, featuring 35.1 billion parameters. It utilizes a unique routed-expert quantization scheme, averaging 2.0586 bits per expert, with every other weight stored in BF16. This model is designed for efficient inference by storing weights dequantized in BF16 tensors, ensuring compatibility with standard Hugging Face Transformers and vLLM libraries.
Loading preview...
Overview
minjaechoi/qwen3p6-35b-a3b-2p06bit-r59 is an internal research checkpoint derived from the Qwen3.6-35B-A3B base model. This 35.1 billion parameter model incorporates a novel quantization approach, where routed experts average 2.0586 bits, while all other weights are maintained in BF16 format. The model's design prioritizes efficient loading and compatibility, as weights are stored dequantized in BF16 tensors, allowing seamless integration with standard transformers and vLLM libraries.
Key Characteristics
- Base Architecture: Built upon the robust Qwen3.6-35B-A3B model.
- Quantization Scheme: Employs a unique routed-expert quantization, achieving an average of 2.0586 bits per expert.
- Hybrid Precision: Combines 2.0586-bit routed experts with BF16 for other weights, balancing efficiency and performance.
- Compatibility: Designed to load directly with stock
transformersand vLLM, simplifying deployment.
Good for
- Research and Experimentation: Ideal for exploring advanced quantization techniques and their impact on large language models.
- Efficient Inference: Suitable for scenarios requiring reduced memory footprint and potentially faster inference due to its quantized structure, while maintaining compatibility with existing ML frameworks.