minjaechoi/qwen3p6-35b-a3b-tileq-2p32bit-r42
minjaechoi/qwen3p6-35b-a3b-tileq-2p32bit-r42 is a 35.1 billion parameter language model based on the Qwen3.6-35B-A3B architecture. This model features an internal research checkpoint utilizing routed experts, where the average weight precision is 2.3203 bits, with other weights stored in BF16. It is designed for efficient loading with standard transformers and vLLM libraries, offering a specialized approach to model quantization.
Loading preview...
Model Overview
This model, minjaechoi/qwen3p6-35b-a3b-tileq-2p32bit-r42, is an internal research checkpoint derived from the Qwen/Qwen3.6-35B-A3B base model. It incorporates a unique quantization strategy focused on "routed experts."
Key Characteristics
- Base Model: Qwen/Qwen3.6-35B-A3B, a 35.1 billion parameter model.
- Quantization: Employs routed experts with an average precision of 2.3203 bits. This means that while some weights are highly quantized, others remain in BF16 format.
- Efficiency: Weights are stored dequantized in BF16 tensors, ensuring compatibility and efficient loading with standard libraries like
transformersandvLLM. - Internal ID: Identified internally as
r42, indicating its status as a specific research iteration.
Potential Use Cases
This model is particularly relevant for researchers and developers interested in:
- Exploring advanced quantization techniques, specifically routed experts.
- Evaluating the performance and efficiency of mixed-precision models.
- Working with Qwen-based architectures that prioritize memory and computational optimization through novel quantization methods.