minjaechoi/qwen3p6-35b-a3b-2p02bit-r71
minjaechoi/qwen3p6-35b-a3b-2p02bit-r71 is a 35.1 billion parameter language model based on the Qwen3.6-35B-A3B architecture. This model features an internal research checkpoint with routed experts averaging 2.0218 bits, while other weights are in BF16 format. It is designed for efficient deployment, loading dequantized weights with standard `transformers` and vLLM libraries. This model is optimized for scenarios requiring a balance of performance and reduced memory footprint through its unique quantization approach.
Loading preview...
Model Overview
minjaechoi/qwen3p6-35b-a3b-2p02bit-r71 is an internal research checkpoint derived from the Qwen/Qwen3.6-35B-A3B base model. This 35.1 billion parameter language model incorporates a unique quantization strategy to optimize its footprint and performance.
Key Characteristics
- Base Architecture: Built upon the robust Qwen3.6-35B-A3B model.
- Quantization: Utilizes routed experts with an average of 2.0218 bits, significantly reducing the precision of specific components.
- Mixed Precision: While routed experts are quantized, all other weights are maintained in BF16 (BFloat16) format, balancing efficiency with numerical stability.
- Deployment Compatibility: Designed for seamless integration with standard deep learning frameworks, loading dequantized weights directly using
transformersand vLLM libraries. - Context Length: Supports a substantial context window of 32768 tokens.
Use Cases
This model is particularly well-suited for developers and researchers looking to:
- Experiment with advanced quantization techniques for large language models.
- Deploy a 35B-class model with reduced memory requirements due to its mixed-precision quantization.
- Leverage the capabilities of the Qwen3.6-35B-A3B architecture in resource-constrained environments.
- Conduct research into the impact and effectiveness of routed expert quantization.