minjaechoi/qwen3p6-35b-a3b-2p03bit-r76

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:2Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 7, 2026Architecture:Transformer Featherless Exclusive Cold

minjaechoi/qwen3p6-35b-a3b-2p03bit-r76 is a 35.1 billion parameter language model based on Qwen/Qwen3.6-35B-A3B, featuring an average of 2.0342-bit routed experts. This model utilizes a hybrid weight storage where routed experts are compressed to approximately 2 bits, while other weights remain in BF16 format. It is designed for efficient loading with standard Hugging Face Transformers and vLLM, making it suitable for research into quantized expert models.

Loading preview...

Overview

minjaechoi/qwen3p6-35b-a3b-2p03bit-r76 is an internal research checkpoint derived from the Qwen/Qwen3.6-35B-A3B base model. This 35.1 billion parameter model explores a novel quantization strategy, where its 'routed experts' are compressed to an average of 2.0342 bits per weight. All other weights within the model are maintained in BF16 precision.

Key Characteristics

  • Hybrid Quantization: Combines highly compressed 2-bit routed experts with standard BF16 weights for other parts of the model.
  • Efficient Loading: Designed to be compatible with stock transformers and vLLM libraries, allowing for straightforward integration and deployment despite its mixed precision.
  • Research Focus: Represents an internal research checkpoint (r76) aimed at investigating the performance and efficiency of models with ultra-low-bit routed experts.

Good for

  • Quantization Research: Ideal for researchers and developers interested in exploring advanced quantization techniques, particularly those involving sparse or expert-based model architectures.
  • Memory-Constrained Inference: Potentially offers advantages in scenarios where memory footprint is a critical concern, due to the significant compression of expert weights.
  • Performance Evaluation: Useful for evaluating the trade-offs between model size, inference speed, and performance when applying aggressive quantization to specific model components.