minjaechoi/qwen3p6-35b-a3b-2p01bit-r61

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:2Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 5, 2026Architecture:Transformer Featherless Exclusive Cold

The minjaechoi/qwen3p6-35b-a3b-2p01bit-r61 is a 35.1 billion parameter language model based on the Qwen3.6-35B-A3B architecture, featuring a context length of 32768 tokens. This model utilizes a unique routing mechanism where experts average 2.0127 bits, with other weights stored in BF16 format. It is designed for efficient loading with standard `transformers` and vLLM libraries, making it suitable for applications requiring a balance of performance and reduced memory footprint through quantized expert routing.

Loading preview...

Overview

The minjaechoi/qwen3p6-35b-a3b-2p01bit-r61 is an internal research checkpoint derived from the Qwen/Qwen3.6-35B-A3B base model. This 35.1 billion parameter language model incorporates a novel approach to quantization, specifically targeting its expert layers.

Key Capabilities

  • Quantized Expert Routing: The model employs a system where its routed experts achieve an average quantization of 2.0127 bits. This significantly reduces the memory footprint and computational requirements for these specific components.
  • Mixed Precision Weights: While experts are heavily quantized, all other weights within the model are maintained in BF16 (BFloat16) format, ensuring a balance between efficiency and numerical precision.
  • Standard Compatibility: Despite its specialized quantization, the model is designed to load seamlessly with conventional deep learning libraries such as transformers and vLLM, as its weights are stored dequantized in BF16 tensors.
  • Large Context Window: Inheriting from its base, the model supports a substantial context length of 32768 tokens, enabling it to process and generate longer sequences of text.

Good for

  • Memory-constrained deployments: Ideal for scenarios where reducing the model's memory footprint is critical, thanks to its highly quantized expert layers.
  • Research into efficient model architectures: Provides a practical example of applying routed expert quantization for performance optimization.
  • Applications requiring a balance of speed and accuracy: The mixed-precision approach aims to offer benefits of quantization without severely compromising the model's overall performance.

This model's license follows that of its base model, Qwen/Qwen3.6-35B-A3B.