minjaechoi/qwen36-35b-a3b-1p80bit-r10

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:2Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 17, 2026Architecture:Transformer Featherless Exclusive Cold

The minjaechoi/qwen36-35b-a3b-1p80bit-r10 model is a 35.1 billion parameter language model based on the Qwen3.6-35B-A3B architecture. This model features an internal research checkpoint utilizing routed experts, where the average expert weight is quantized to 1.80 bits, while other weights remain in BF16. It is designed for efficient inference by storing weights dequantized in BF16 tensors, compatible with stock `transformers` and vLLM.

Loading preview...

Model Overview

The minjaechoi/qwen36-35b-a3b-1p80bit-r10 is an internal research checkpoint derived from the Qwen/Qwen3.6-35B-A3B base model. This 35.1 billion parameter model incorporates a unique quantization strategy to optimize efficiency.

Key Characteristics

  • Quantization: Employs routed experts with an average quantization of 1.80 bits per expert. All other weights are maintained in BF16 format.
  • Compatibility: Weights are stored dequantized in BF16 tensors, ensuring seamless loading and operation with standard libraries like transformers and vLLM.
  • Base Model: Built upon the robust Qwen3.6-35B-A3B architecture, inheriting its foundational capabilities.

Intended Use Cases

This model is particularly suited for research and development environments focused on:

  • Efficient Inference: Leveraging the 1.80-bit routed experts for potentially faster and more memory-efficient operations compared to full BF16 models.
  • Exploration of Quantization: Investigating the performance and trade-offs of advanced quantization techniques in large language models.
  • Compatibility Testing: Utilizing its stock transformers and vLLM compatibility for integration into existing inference pipelines.