minjaechoi/qwen3p6-35b-a3b-2p04bit-r60
minjaechoi/qwen3p6-35b-a3b-2p04bit-r60 is an internal research checkpoint based on the Qwen3.6-35B-A3B model, featuring 35.1 billion parameters and a 32768-token context length. This model utilizes routed experts with an average of 2.0431 bits per expert, while other weights are in BF16 format. It is designed for research into efficient model architectures, allowing for loading with standard Hugging Face Transformers and vLLM libraries.
Loading preview...
Model Overview
minjaechoi/qwen3p6-35b-a3b-2p04bit-r60 is an internal research checkpoint derived from the Qwen3.6-35B-A3B base model. This 35.1 billion parameter model incorporates a unique architecture where routed experts average 2.0431 bits for their weights, while all other weights are maintained in BF16 format. It supports a substantial context length of 32768 tokens.
Key Characteristics
- Efficient Weight Representation: Employs a hybrid quantization scheme with ultra-low-bit routed experts (averaging 2.0431 bits) alongside BF16 weights for other parameters.
- Standard Compatibility: Despite its specialized weight representation, the model is designed to be loaded and utilized with conventional
transformersandvLLMlibraries, as weights are stored dequantized in BF16 tensors. - Research Focus: Primarily intended as an internal research checkpoint (identified as
r60) for exploring and developing efficient large language model architectures.
Good for
- Advanced LLM Research: Ideal for researchers investigating novel quantization techniques, routed expert models, and efficient inference strategies.
- Performance Optimization Studies: Suitable for experiments aimed at reducing model size and improving inference speed while maintaining performance.
- Exploring Qwen3.6-35B-A3B Variants: Provides a specialized version of the Qwen3.6-35B-A3B base model for comparative analysis of different architectural and quantization approaches.