minjaechoi/qwen3p6-35b-a3b-2p00bit-r52

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:2Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 28, 2026Architecture:Transformer Featherless Exclusive Cold

The minjaechoi/qwen3p6-35b-a3b-2p00bit-r52 model is an internal research checkpoint based on the Qwen3.6-35B-A3B architecture, featuring 35.1 billion parameters and a 32768-token context length. This model utilizes routed experts with an average quantization of 2.0000 bits, while other weights remain in BF16 format. It is designed for efficient inference with quantized weights that load directly using standard `transformers` and vLLM libraries.

Loading preview...

Model Overview

This model, minjaechoi/qwen3p6-35b-a3b-2p00bit-r52, is an internal research checkpoint derived from the Qwen/Qwen3.6-35B-A3B base model. It incorporates a unique quantization strategy to optimize performance and efficiency.

Key Characteristics

  • Base Architecture: Qwen3.6-35B-A3B, a large language model with 35.1 billion parameters.
  • Quantization: Employs routed experts with an average quantization of 2.0000 bits. All other weights are maintained in BF16 (Brain Floating Point) format.
  • Loading Compatibility: Weights are stored dequantized in BF16 tensors, ensuring seamless loading and compatibility with standard libraries such as transformers and vLLM.
  • Context Length: Supports a substantial context window of 32768 tokens, suitable for processing longer inputs.

Use Cases

This model is particularly suited for research and development environments where the goal is to explore the trade-offs between model size, performance, and computational efficiency through advanced quantization techniques. Its compatibility with common LLM inference frameworks makes it accessible for experimentation with quantized models.