minjaechoi/qwen3p6-35b-a3b-nopin-2p00bit-r40
minjaechoi/qwen3p6-35b-a3b-nopin-2p00bit-r40 is a 35.1 billion parameter language model based on the Qwen3.6-35B-A3B architecture. This model features an internal research checkpoint utilizing routed experts quantized to an average of 2.0000 bits, while other weights remain in BF16. It is designed to load with standard `transformers` and vLLM libraries, offering a balance of performance and reduced memory footprint through its unique quantization strategy.
Loading preview...
Model Overview
This model, minjaechoi/qwen3p6-35b-a3b-nopin-2p00bit-r40, is an internal research checkpoint derived from the Qwen3.6-35B-A3B base model. It features a 35.1 billion parameter architecture with a notable quantization strategy.
Key Characteristics
- Base Architecture: Built upon the robust Qwen3.6-35B-A3B model, known for its general language understanding and generation capabilities.
- Quantization: Employs a unique routing expert quantization where experts average 2.0000 bits, significantly reducing their memory footprint. All other weights are maintained in BF16 precision.
- Memory Efficiency: The routed expert quantization allows for a more memory-efficient model while aiming to preserve performance.
- Compatibility: Designed for seamless integration with standard machine learning libraries, loading directly with
transformersand vLLM without requiring specialized loaders for the quantized components. - License: Adheres to the licensing terms of its base model, Qwen/Qwen3.6-35B-A3B.
Potential Use Cases
This model is particularly interesting for developers and researchers focused on:
- Resource-constrained deployments: Its quantization strategy makes it suitable for environments where memory or computational resources are limited.
- Experimentation with quantization techniques: Offers a practical example of routed expert quantization for further research and development.
- Applications requiring a powerful 35B-class model: When a balance between model size and efficiency is crucial, leveraging the Qwen3.6-35B-A3B's capabilities with reduced bit-width experts.