minjaechoi/qwen3p6-35b-a3b-2p05bit-r67
minjaechoi/qwen3p6-35b-a3b-2p05bit-r67 is a 35.1 billion parameter language model based on Qwen3.6-35B-A3B, featuring an average of 2.0460-bit routed experts. This internal research checkpoint stores weights dequantized in BF16 tensors, ensuring compatibility with standard Hugging Face Transformers and vLLM. Its primary differentiator is the highly compressed, routed expert architecture, making it suitable for efficient deployment while maintaining performance.
Loading preview...
Overview
minjaechoi/qwen3p6-35b-a3b-2p05bit-r67 is an internal research checkpoint derived from the Qwen3.6-35B-A3B base model. This 35.1 billion parameter language model incorporates a unique routed expert architecture, achieving an average of 2.0460 bits per expert. The model's weights are stored dequantized in BF16 tensors, allowing for seamless integration and loading with standard transformers and vLLM libraries.
Key Capabilities
- Highly Efficient Compression: Utilizes a routed expert architecture to achieve an average of 2.0460 bits, significantly reducing model size and computational requirements.
- Standard Library Compatibility: Designed to load effortlessly with existing
transformersandvLLMframeworks, simplifying deployment. - Based on Qwen3.6-35B-A3B: Inherits the foundational capabilities and performance characteristics of its robust base model.
Good For
- Resource-Constrained Deployments: Ideal for applications requiring a powerful language model with a reduced memory footprint and faster inference.
- Research and Experimentation: Suitable for exploring the performance and efficiency benefits of highly compressed, routed expert architectures.
- Edge Device Inference: Potentially beneficial for running large language models on devices with limited computational resources.