minjaechoi/qwen3p6-35b-a3b-2p06bit-r58
TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:2Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 5, 2026Architecture:Transformer Featherless Exclusive Cold
The minjaechoi/qwen3p6-35b-a3b-2p06bit-r58 model is an internal research checkpoint based on the Qwen/Qwen3.6-35B-A3B architecture, featuring 35.1 billion parameters and a 32768-token context length. This model utilizes a unique quantization scheme where routed experts average 2.0578 bits, while other weights remain in BF16 format. It is designed for efficient loading with standard `transformers` and vLLM libraries, making it suitable for research into highly quantized large language models.
Loading preview...
Model Overview
This model, minjaechoi/qwen3p6-35b-a3b-2p06bit-r58, is an internal research checkpoint derived from the Qwen/Qwen3.6-35B-A3B base model. It features a substantial 35.1 billion parameters and supports a context length of 32768 tokens.
Key Characteristics
- Quantization Scheme: A primary differentiator is its advanced quantization, where "routed experts" are compressed to an average of 2.0578 bits. All other weights are maintained in BF16 precision.
- Efficiency: Despite the quantization, weights are stored dequantized in BF16 tensors, ensuring compatibility and efficient loading with standard libraries like
transformersand vLLM. - Research Focus: This model represents an internal research effort, likely exploring the performance and efficiency trade-offs of highly quantized large language models.
Potential Use Cases
- Quantization Research: Ideal for researchers investigating extreme quantization techniques and their impact on LLM performance and inference costs.
- Efficient Deployment: Could serve as a foundation for applications requiring a large language model with reduced memory footprint and potentially faster inference, provided the quantization trade-offs are acceptable for the specific task.
- Experimental Applications: Suitable for developers and researchers looking to experiment with cutting-edge model compression methods.