minjaechoi/qwen36-35b-a3b-2p00bit-r7

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:2Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 17, 2026Architecture:Transformer Featherless Exclusive Cold

The minjaechoi/qwen36-35b-a3b-2p00bit-r7 model is a 35.1 billion parameter language model based on the Qwen3.6-35B-A3B architecture. This version features an internal research checkpoint utilizing routed experts with an average of 2.00 bits, where every other weight is in BF16 format. Weights are stored dequantized in BF16 tensors, ensuring compatibility with standard `transformers` and vLLM libraries. It is designed for efficient deployment while maintaining performance characteristics of its base model.

Loading preview...

Model Overview

minjaechoi/qwen36-35b-a3b-2p00bit-r7 is a 35.1 billion parameter language model derived from the Qwen/Qwen3.6-35B-A3B base model. This particular iteration represents an internal research checkpoint, identified as 'r7', focusing on optimized weight representation.

Key Characteristics

  • Base Architecture: Built upon the robust Qwen3.6-35B-A3B model.
  • Quantization Scheme: Employs a unique quantization approach with "routed experts" averaging 2.00 bits. This means that while some weights are highly compressed, every other weight is maintained in BF16 precision.
  • Deployment Compatibility: Despite its specialized quantization, the model's weights are stored dequantized in BF16 tensors, allowing for seamless loading and inference using standard libraries like transformers and vLLM without requiring custom loaders.
  • Context Length: Supports a substantial context length of 32768 tokens, enabling processing of longer inputs and generating more coherent, extended outputs.

Potential Use Cases

This model is suitable for applications requiring a powerful language model with a large context window, where computational efficiency and reduced memory footprint are beneficial. Its compatibility with standard inference frameworks makes it a practical choice for:

  • Text Generation: Creating long-form content, articles, or creative writing.
  • Complex Reasoning: Handling tasks that benefit from extensive contextual understanding.
  • Summarization: Condensing large documents or conversations.
  • Efficient Deployment: Projects where the balance between performance and resource utilization is critical, leveraging its optimized weight representation.