minjaechoi/qwen3p6-35b-a3b-2p02bit-r70

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:2Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 6, 2026Architecture:Transformer Featherless Exclusive Cold

The minjaechoi/qwen3p6-35b-a3b-2p02bit-r70 is a 35.1 billion parameter language model based on the Qwen3.6-35B-A3B architecture. This model features an internal research checkpoint utilizing routed experts, achieving an average of 2.0218 bits per expert while other weights remain in BF16 format. Its primary differentiator is the efficient storage of weights, which are dequantized into BF16 tensors and are compatible with standard transformers and vLLM loading mechanisms. This model is designed for research into efficient model quantization and routing.

Loading preview...

Overview

The minjaechoi/qwen3p6-35b-a3b-2p02bit-r70 is an internal research checkpoint derived from the Qwen/Qwen3.6-35B-A3B base model. This 35.1 billion parameter model explores an advanced quantization technique involving routed experts, which average 2.0218 bits per expert. All other weights within the model are maintained in BF16 format.

Key Characteristics

  • Base Architecture: Built upon the robust Qwen/Qwen3.6-35B-A3B model.
  • Quantization Innovation: Implements a novel routed-expert quantization scheme, achieving a highly efficient average of 2.0218 bits for expert weights.
  • Hybrid Precision: Combines ultra-low-bit routed experts with BF16 precision for other model weights.
  • Compatibility: Weights are stored dequantized in BF16 tensors, ensuring seamless loading with standard transformers and vLLM libraries.
  • Research Focus: Represents an internal research checkpoint (internal ID: r70) aimed at advancing efficient model deployment and performance.

Good For

  • Research and Development: Ideal for researchers exploring advanced quantization techniques, particularly routed experts and hybrid precision methods.
  • Efficiency Studies: Suitable for evaluating the trade-offs between model size, performance, and computational efficiency using novel quantization approaches.
  • Custom Deployments: Potentially useful for scenarios requiring highly optimized models that can still leverage standard inference frameworks.