MingZwhy/Qwen3-1.7B-W1.88-QAD
MingZwhy/Qwen3-1.7B-W1.88-QAD is a 2 billion parameter quantization-aware-distillation checkpoint for the Qwen3-1.7B model, specifically for the W1.88 quantization arm. Developed by MingZwhy, this model serves as a starting point for on-policy distillation training, not a ready-to-use quantized model. It features mixed INT1.58/INT4 weights, INT4 embeddings and output head, and INT8 activations, designed for further quantization-aware training workflows.
Loading preview...
Overview
This model, MingZwhy/Qwen3-1.7B-W1.88-QAD, is a quantization-aware-distillation (QAD) checkpoint for the Qwen3-1.7B architecture, specifically configured for the W1.88 quantization arm. Developed by MingZwhy, it represents a crucial starting point for training in an on-policy distillation (OPD) workflow, rather than a directly deployable, finished model.
Key Characteristics
- Latent Checkpoint: The saved tensors are in
bf16format and are not yet quantized. They represent the high-precision master weights used during quantization-aware training. - Quantization Scheme: Designed for a target quantization of:
- Weights: Mixed INT1.58 / INT4 in blocks of 256, with an effective average of 1.88 bits.
- Embedding & Output Head: INT4.
- Activations: INT8.
- KV Cache: 16-bit during OPD and evaluation.
- Intended Use: This checkpoint is meant to be used as the
STUDENT_MODELin the OPD stage, where the quantizer configuration is applied during training.
Important Note for Users
- Loading this checkpoint directly will result in an unquantized model, which will perform better than the intended W1.88 quantized model.
- For a ready-to-load and evaluate quantized model, users should refer to the recovered checkpoint: MingZwhy/Qwen3-1.7B-W1.88-QAOPD.
- The associated code and recipe for its use in quantization-aware on-policy distillation are available at MingZwhy/QAOPD.