minjaechoi/qwen36-35b-a3b-1p80bit-r11

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:2Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 17, 2026Architecture:Transformer Featherless Exclusive Cold

The minjaechoi/qwen36-35b-a3b-1p80bit-r11 model is a 35.1 billion parameter language model based on the Qwen3.6-35B-A3B architecture. Developed by minjaechoi, this model features routed experts averaging 1.80 bits, with other weights stored in BF16 tensors. It is an internal research checkpoint focused on efficient weight representation, making it suitable for applications requiring reduced memory footprint while maintaining performance.

Loading preview...

Model Overview

The minjaechoi/qwen36-35b-a3b-1p80bit-r11 is an internal research checkpoint derived from the Qwen/Qwen3.6-35B-A3B base model. This 35.1 billion parameter language model incorporates a unique weight representation strategy, utilizing "routed experts" that average 1.80 bits. All other weights are stored in BF16 tensors, which are dequantized and loaded using standard transformers and vLLM libraries.

Key Characteristics

  • Base Architecture: Built upon the robust Qwen3.6-35B-A3B model.
  • Efficient Weight Representation: Employs routed experts with an average of 1.80 bits, alongside BF16 weights, aiming for memory efficiency.
  • Standard Compatibility: Weights are designed to load seamlessly with stock transformers and vLLM, ensuring ease of integration.
  • Research Focus: Represents an internal research checkpoint, indicating exploration into advanced quantization and expert routing techniques.

Potential Use Cases

This model is particularly interesting for:

  • Memory-constrained deployments: Its efficient weight representation could be beneficial for environments where memory footprint is a critical concern.
  • Research into quantization and sparse models: Developers and researchers exploring advanced quantization methods or routed expert architectures may find this model a valuable reference.
  • Applications requiring Qwen3.6-35B-A3B capabilities with efficiency: For tasks where the performance of the base Qwen model is desired, but with an optimized memory profile.