minjaechoi/qwen3p6-35b-a3b-2p09bit-r66

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:2Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 7, 2026Architecture:Transformer Featherless Exclusive Cold

minjaechoi/qwen3p6-35b-a3b-2p09bit-r66 is a 35.1 billion parameter language model based on the Qwen3.6-35B-A3B architecture. This model features routed experts with an average quantization of 2.0920 bits, while other weights remain in BF16 format. It is an internal research checkpoint designed for efficient inference by storing weights dequantized in BF16 tensors, compatible with standard `transformers` and vLLM loading. This model is optimized for scenarios requiring a balance between performance and reduced memory footprint through advanced quantization techniques.

Loading preview...

Model Overview

minjaechoi/qwen3p6-35b-a3b-2p09bit-r66 is an experimental 35.1 billion parameter language model derived from the Qwen/Qwen3.6-35B-A3B base model. This version incorporates a unique quantization strategy, utilizing routed experts that achieve an average of 2.0920 bits per weight. All other weights within the model are maintained in BF16 precision.

Key Technical Details

  • Base Model: Qwen/Qwen3.6-35B-A3B
  • Quantization: Employs routed experts with an average bit-width of 2.0920 bits.
  • Weight Storage: Weights are stored dequantized in BF16 tensors, ensuring compatibility with standard inference libraries like transformers and vLLM.
  • Nature: This is an internal research checkpoint, identified internally as r66.

Intended Use and Benefits

This model is designed for researchers and developers exploring advanced quantization methods to optimize large language models. Its primary benefit lies in demonstrating a practical application of mixed-precision quantization, where specific parts of the model (routed experts) are aggressively quantized while maintaining BF16 for other components. This approach aims to reduce memory footprint and potentially improve inference speed without significant performance degradation, making it suitable for environments with memory constraints or for exploring efficient LLM deployment strategies. The license follows that of the base Qwen model.