MingZwhy/Qwen3-4B-W2.79-QAD

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 21, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

MingZwhy/Qwen3-4B-W2.79-QAD is a 4 billion parameter Qwen3-based model developed by MingZwhy, serving as a quantization-aware-distillation (QAD) checkpoint. This model is specifically designed as a starting point for on-policy distillation (OPD) training, targeting efficient deployment with a W2.79 effective bitwidth for weights. It features mixed INT1.58/INT4 weights, INT4 for embedding and output head, and INT8 activations, making it optimized for further quantization-aware training rather than direct inference.

Loading preview...

MingZwhy/Qwen3-4B-W2.79-QAD: A Quantization-Aware-Distillation Checkpoint

This model, developed by MingZwhy, is a 4 billion parameter Qwen3-based checkpoint specifically engineered for quantization-aware-distillation (QAD). It represents the initial state for on-policy distillation (OPD) within the W2.79 bitwidth arm, making it a crucial starting point for training rather than a finished, directly deployable model.

Key Characteristics

  • Latent Checkpoint: The tensors are in bf16 format and are not yet quantized. Loading this file directly provides an unquantized model, which will perform better than the intended W2.79 quantized version until further training.
  • Quantization Scheme: Designed for a target effective bitwidth of 2.79 bits for weights. This is achieved through a mixed INT1.58 / INT4 scheme in blocks of 256, with 50% of blocks at INT4.
  • Component Quantization: Embedding and output head are INT4, while activations are INT8. The KV cache maintains 16-bit precision during OPD and evaluation.

Intended Use

This checkpoint is primarily intended to be used as the STUDENT_MODEL for the OPD stage, where the quantizer configuration is applied during training. For a directly loadable and evaluable quantized model, users should refer to MingZwhy/Qwen3-4B-W2.79-QAOPD instead. The associated code and recipe for its use are available at MingZwhy/QAOPD.