minjaechoi/qwen36-35b-a3b-1p80bit-r10
The minjaechoi/qwen36-35b-a3b-1p80bit-r10 model is a 35.1 billion parameter language model based on the Qwen3.6-35B-A3B architecture. This model features an internal research checkpoint utilizing routed experts, where the average expert weight is quantized to 1.80 bits, while other weights remain in BF16. It is designed for efficient inference by storing weights dequantized in BF16 tensors, compatible with stock `transformers` and vLLM.
Loading preview...
Model Overview
The minjaechoi/qwen36-35b-a3b-1p80bit-r10 is an internal research checkpoint derived from the Qwen/Qwen3.6-35B-A3B base model. This 35.1 billion parameter model incorporates a unique quantization strategy to optimize efficiency.
Key Characteristics
- Quantization: Employs routed experts with an average quantization of 1.80 bits per expert. All other weights are maintained in BF16 format.
- Compatibility: Weights are stored dequantized in BF16 tensors, ensuring seamless loading and operation with standard libraries like
transformersand vLLM. - Base Model: Built upon the robust Qwen3.6-35B-A3B architecture, inheriting its foundational capabilities.
Intended Use Cases
This model is particularly suited for research and development environments focused on:
- Efficient Inference: Leveraging the 1.80-bit routed experts for potentially faster and more memory-efficient operations compared to full BF16 models.
- Exploration of Quantization: Investigating the performance and trade-offs of advanced quantization techniques in large language models.
- Compatibility Testing: Utilizing its stock
transformersand vLLM compatibility for integration into existing inference pipelines.