minjaechoi/qwen3p6-35b-a3b-2p32bit-r51

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:2Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 28, 2026Architecture:Transformer Featherless Exclusive Cold

The minjaechoi/qwen3p6-35b-a3b-2p32bit-r51 is a 35.1 billion parameter language model based on the Qwen3.6-35B-A3B architecture, featuring an internal research checkpoint with routed experts. This model utilizes an average of 2.3219 bits for its routed experts, while other weights are in BF16 format. It is designed for efficient loading with standard transformers and vLLM libraries, making it suitable for applications requiring a balance of performance and reduced memory footprint.

Loading preview...

Model Overview

The minjaechoi/qwen3p6-35b-a3b-2p32bit-r51 is an internal research checkpoint derived from the Qwen/Qwen3.6-35B-A3B base model. This 35.1 billion parameter model incorporates a unique quantization strategy to optimize its footprint and performance.

Key Characteristics

  • Quantization Scheme: Features routed experts with an average of 2.3219 bits, significantly reducing the precision for these components.
  • Mixed Precision: While routed experts are highly quantized, all other weights are maintained in BF16 (BFloat16) format, balancing efficiency with numerical stability.
  • Compatibility: Weights are stored dequantized in BF16 tensors, ensuring seamless loading and operation with standard transformers and vLLM libraries.
  • Base Model: Built upon the robust Qwen3.6-35B-A3B architecture, inheriting its foundational capabilities.

Intended Use Cases

This model is particularly well-suited for developers and researchers exploring:

  • Efficient Deployment: Its mixed-precision quantization makes it a strong candidate for environments where memory and computational resources are constrained, without sacrificing too much performance.
  • Research in Quantization: Provides a practical example of routed expert quantization, offering insights into advanced model compression techniques.
  • Applications requiring Qwen3.6-35B-A3B capabilities: Can be used in scenarios where the base model's strengths are desired, but with improved inference efficiency due to the quantization.