minjaechoi/qwen3p6-35b-a3b-2p05bit-r62

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:2Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 5, 2026Architecture:Transformer Featherless Exclusive Cold

minjaechoi/qwen3p6-35b-a3b-2p05bit-r62 is a 35.1 billion parameter language model based on Qwen/Qwen3.6-35B-A3B, featuring a unique routed-expert architecture. This model utilizes an average of 2.0538 bits per routed expert, with every other weight stored in BF16 format. It is designed for efficient loading with standard `transformers` and vLLM libraries, making it suitable for research into quantized expert models.

Loading preview...

Overview

minjaechoi/qwen3p6-35b-a3b-2p05bit-r62 is an internal research checkpoint derived from the Qwen/Qwen3.6-35B-A3B base model. This 35.1 billion parameter model incorporates a novel routed-expert architecture, distinguishing it from standard large language models.

Key Characteristics

  • Quantized Routed Experts: The model employs routed experts with an average quantization of 2.0538 bits. This significantly reduces the memory footprint and computational requirements for the expert layers.
  • Hybrid Precision: While experts are quantized, every other weight in the model is maintained in BF16 (bfloat16) precision, balancing efficiency with model accuracy.
  • Standard Compatibility: Despite its unique quantization scheme, the weights are stored dequantized in BF16 tensors, allowing for seamless loading and inference using standard libraries like transformers and vLLM.
  • Research Focus: Identified as an "internal research checkpoint" (r62), this model is primarily intended for advanced research into efficient model architectures and quantization techniques.

Good for

  • Research into Quantization: Ideal for researchers exploring advanced quantization methods, particularly routed-expert architectures and hybrid precision schemes.
  • Efficient Inference Experimentation: Suitable for experimenting with LLMs that aim to reduce memory and computational overhead while maintaining performance.
  • Understanding Qwen Architectures: Provides insights into potential future optimizations for Qwen-based models.