MingZwhy/Qwen3-4B-W1.88-QAD
MingZwhy/Qwen3-4B-W1.88-QAD is a 4 billion parameter quantization-aware-distillation (QAD) checkpoint for the Qwen3 model family, developed by MingZwhy. This model serves as a starting point for on-policy distillation (OPD) training, specifically for the W1.88 quantization arm. It is designed to be used as a student model in further quantization training processes, rather than a finished, directly loadable model for inference. Its primary purpose is to facilitate the development of highly quantized models with an effective weight bit-width of 1.88 bits.
Loading preview...
MingZwhy/Qwen3-4B-W1.88-QAD: A Quantization-Aware-Distillation Checkpoint
This model, developed by MingZwhy, is a quantization-aware-distillation (QAD) checkpoint for the Qwen3-4B architecture, specifically configured for a W1.88 quantization scheme. It represents the initial state from which on-policy distillation (OPD) training commences.
Key Characteristics & Usage
- Starting Point for Training: This checkpoint is not a ready-to-use model for direct inference. Instead, it's intended as a
STUDENT_MODELfor the OPD stage, where the quantizer configuration is applied during training. - Latent Checkpoint: The tensors are in
bf16format and are not yet quantized. Quantization-aware training maintains high-precision master weights, with quantization applied within the forward pass. Loading this file directly yields an unquantized model. - Quantization Scheme: When used in the intended OPD process, it targets a mixed INT1.58 / INT4 weight quantization in blocks, resulting in an effective 1.88 bits for weights. Embedding and output heads are INT4, and activations are INT8.
- Recommended Alternative: For a directly loadable and evaluable quantized model, users should refer to MingZwhy/Qwen3-4B-W1.88-QAOPD instead.
Intended Use Case
This model is specifically designed for researchers and developers engaged in quantization research and development, particularly those working on on-policy distillation to achieve highly efficient, low-bit models based on the Qwen3-4B architecture. It provides the foundational checkpoint for initiating such advanced quantization training workflows.