minjaechoi/qwen3p6-35b-a3b-2p02bit-r72
The minjaechoi/qwen3p6-35b-a3b-2p02bit-r72 is a 35.1 billion parameter language model based on the Qwen3.6-35B-A3B architecture. This model features routed experts with an average quantization of 2.0218 bits, while other weights remain in BF16 format. It is designed for efficient loading with standard `transformers` and vLLM libraries, making it suitable for applications requiring a balance of performance and reduced memory footprint.
Loading preview...
Model Overview
The minjaechoi/qwen3p6-35b-a3b-2p02bit-r72 is an internal research checkpoint derived from the Qwen/Qwen3.6-35B-A3B base model. This 35.1 billion parameter model incorporates a unique quantization strategy to optimize its footprint and performance.
Key Characteristics
- Base Architecture: Built upon the robust Qwen3.6-35B-A3B model.
- Quantization: Utilizes a routed-expert approach where experts average 2.0218 bits. All other weights are maintained in BF16 precision.
- Efficiency: Weights are stored dequantized in BF16 tensors, ensuring compatibility and efficient loading with standard libraries like
transformersand vLLM. - Context Length: Supports a context length of 32768 tokens, enabling processing of longer sequences.
Use Cases
This model is particularly suited for developers and researchers looking for:
- Memory-efficient deployments: The mixed-precision quantization helps reduce memory requirements compared to full BF16 models of similar size.
- Research into quantization techniques: Provides a practical example of routed-expert quantization in a large language model.
- Applications requiring Qwen3.6 capabilities: Leverages the underlying strengths of the Qwen3.6-35B-A3B base model while offering efficiency improvements.
The license for this model follows that of its base model, Qwen/Qwen3.6-35B-A3B.