minjaechoi/qwen3p6-35b-a3b-2p06bit-r59

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:2Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 5, 2026Architecture:Transformer0.0K Featherless Exclusive Cold

The minjaechoi/qwen3p6-35b-a3b-2p06bit-r59 model is an internal research checkpoint based on the Qwen3.6-35B-A3B architecture, featuring 35.1 billion parameters. It utilizes a unique routed-expert quantization scheme, averaging 2.0586 bits per expert, with every other weight stored in BF16. This model is designed for efficient inference by storing weights dequantized in BF16 tensors, ensuring compatibility with standard Hugging Face Transformers and vLLM libraries.

Loading preview...

Overview

minjaechoi/qwen3p6-35b-a3b-2p06bit-r59 is an internal research checkpoint derived from the Qwen3.6-35B-A3B base model. This 35.1 billion parameter model incorporates a novel quantization approach, where routed experts average 2.0586 bits, while all other weights are maintained in BF16 format. The model's design prioritizes efficient loading and compatibility, as weights are stored dequantized in BF16 tensors, allowing seamless integration with standard transformers and vLLM libraries.

Key Characteristics

  • Base Architecture: Built upon the robust Qwen3.6-35B-A3B model.
  • Quantization Scheme: Employs a unique routed-expert quantization, achieving an average of 2.0586 bits per expert.
  • Hybrid Precision: Combines 2.0586-bit routed experts with BF16 for other weights, balancing efficiency and performance.
  • Compatibility: Designed to load directly with stock transformers and vLLM, simplifying deployment.

Good for

  • Research and Experimentation: Ideal for exploring advanced quantization techniques and their impact on large language models.
  • Efficient Inference: Suitable for scenarios requiring reduced memory footprint and potentially faster inference due to its quantized structure, while maintaining compatibility with existing ML frameworks.