minjaechoi/qwen3p6-35b-a3b-2p02bit-r31_v7
minjaechoi/qwen3p6-35b-a3b-2p02bit-r31_v7 is a 35.1 billion parameter language model based on the Qwen3.6-35B-A3B architecture. This model features an internal research checkpoint utilizing routed experts, achieving an average of 2.022 bits per expert. It is designed for efficient inference, with weights stored dequantized in BF16 tensors and compatible with standard `transformers` and vLLM libraries. This version, identified as r31_v7, focuses on exploring quantization techniques for large language models.
Loading preview...
Model Overview
minjaechoi/qwen3p6-35b-a3b-2p02bit-r31_v7 is an internal research checkpoint derived from the Qwen3.6-35B-A3B base model. This 35.1 billion parameter model incorporates a novel quantization approach using routed experts, where the experts average 2.022 bits.
Key Characteristics
- Base Architecture: Built upon the robust Qwen3.6-35B-A3B model.
- Quantization: Employs a routed-expert quantization scheme, achieving an average of 2.022 bits per expert, while other weights remain in BF16 format.
- Efficiency: Weights are stored dequantized in BF16 tensors, ensuring compatibility and efficient loading with standard libraries like
transformersand vLLM. - Research Focus: Represents an internal research checkpoint (r31_v7) exploring advanced quantization and expert routing techniques for large language models.
Potential Use Cases
- Research and Development: Ideal for researchers and developers interested in experimenting with highly quantized models and routed expert architectures.
- Efficient Deployment: Suitable for scenarios requiring reduced memory footprint and potentially faster inference speeds compared to full-precision models of similar size, while maintaining compatibility with existing inference frameworks.