minjaechoi/qwen36-35b-a3b-2p00bit-r8
minjaechoi/qwen36-35b-a3b-2p00bit-r8 is a 35.1 billion parameter language model based on the Qwen3.6-35B-A3B architecture, featuring an internal research checkpoint with routed experts averaging 2.00 bits. This model utilizes BF16 for other weights and stores dequantized weights in BF16 tensors, ensuring compatibility with standard `transformers` and vLLM. It is designed for efficient deployment while maintaining performance, leveraging its unique quantization strategy.
Loading preview...
Model Overview
minjaechoi/qwen36-35b-a3b-2p00bit-r8 is an internal research checkpoint derived from the Qwen3.6-35B-A3B base model. This 35.1 billion parameter model incorporates a specialized quantization technique, where its routed experts average 2.00 bits per weight, while all other weights are maintained in BF16 format.
Key Characteristics
- Quantization Strategy: Features a hybrid quantization approach with 2.00-bit routed experts and BF16 for remaining weights.
- Compatibility: Weights are stored dequantized in BF16 tensors, ensuring seamless loading and operation with standard
transformersand vLLM libraries. - Efficiency Focus: This configuration aims to provide a balance between model performance and computational efficiency, particularly in terms of memory footprint and inference speed.
Good For
- Research and Development: Ideal for researchers exploring advanced quantization techniques and their impact on large language models.
- Efficient Deployment: Suitable for applications requiring a powerful 35B parameter model with optimized memory usage and faster inference due to its unique bit-level routing.
- Qwen3.6-35B-A3B Users: Developers already familiar with the Qwen3.6-35B-A3B base model who are looking for a more resource-efficient variant.