minjaechoi/qwen36-35b-a3b-2p00bit-r13
minjaechoi/qwen36-35b-a3b-2p00bit-r13 is a 35.1 billion parameter language model based on the Qwen3.6-35B-A3B architecture. This model features an internal research checkpoint utilizing routed experts, where the average expert weight is quantized to 2.00 bits, while other weights remain in BF16. It is designed for efficient deployment, loading dequantized weights with standard transformers and vLLM libraries, making it suitable for applications requiring a balance of performance and reduced memory footprint.
Loading preview...
Model Overview
minjaechoi/qwen36-35b-a3b-2p00bit-r13 is an internal research checkpoint derived from the Qwen/Qwen3.6-35B-A3B base model. This 35.1 billion parameter model incorporates a unique quantization strategy focused on routed experts.
Key Characteristics
- Quantization: The model employs a mixed-precision approach where routed experts average 2.00 bits per weight, significantly reducing their memory footprint. All other weights are maintained in BF16 precision.
- Architecture: Built upon the robust Qwen3.6-35B-A3B architecture, known for its strong performance in various language tasks.
- Deployment: Designed for ease of integration, the weights are stored dequantized in BF16 tensors, allowing for seamless loading and inference using standard
transformersandvLLMlibraries. - Efficiency: The 2.00-bit routed experts contribute to a more efficient model, potentially offering faster inference or reduced memory consumption compared to full-precision alternatives of similar scale.
Use Cases
This model is particularly well-suited for developers and researchers looking to:
- Experiment with highly quantized large language models that maintain compatibility with standard inference frameworks.
- Deploy a powerful 35.1B parameter model in environments where memory efficiency is a critical factor.
- Explore the performance characteristics of models utilizing routed experts for conditional computation and sparsity.