minjaechoi/qwen36-35b-a3b-2p00bit-r12
TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:2Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 17, 2026Architecture:Transformer Featherless Exclusive Cold
The minjaechoi/qwen36-35b-a3b-2p00bit-r12 is a 35.1 billion parameter language model based on the Qwen3.6-35B-A3B architecture. This model features routed experts with an average quantization of 2.00 bits, while other weights are stored in BF16. It is an internal research checkpoint designed for efficient inference with reduced memory footprint, loading with standard `transformers` and vLLM libraries.
Loading preview...
Model Overview
The minjaechoi/qwen36-35b-a3b-2p00bit-r12 is an internal research checkpoint derived from the Qwen/Qwen3.6-35B-A3B base model. This 35.1 billion parameter model incorporates a unique quantization strategy to optimize performance and resource usage.
Key Characteristics
- Quantization: Utilizes routed experts with an average quantization of 2.00 bits. This means that while some parts of the model are highly compressed, other weights remain in BF16 format for precision.
- Base Model: Built upon the robust Qwen/Qwen3.6-35B-A3B architecture, inheriting its core capabilities.
- Compatibility: Designed to be loaded and used with standard
transformersand vLLM libraries, ensuring ease of integration into existing workflows. - Memory Efficiency: The mixed-precision approach, combining 2.00-bit routed experts with BF16 weights, aims to reduce the model's memory footprint during inference.
Potential Use Cases
This model is suitable for developers and researchers interested in:
- Efficient Inference: Deploying large language models with reduced memory requirements.
- Quantization Research: Exploring the impact and performance of mixed-precision quantization techniques.
- Resource-Constrained Environments: Utilizing a powerful 35B parameter model where memory optimization is critical.