MingZwhy/Qwen3-4B-W1.88-QAD

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 21, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

MingZwhy/Qwen3-4B-W1.88-QAD is a 4 billion parameter quantization-aware-distillation (QAD) checkpoint for the Qwen3 model family, developed by MingZwhy. This model serves as a starting point for on-policy distillation (OPD) training, specifically for the W1.88 quantization arm. It is designed to be used as a student model in further quantization training processes, rather than a finished, directly loadable model for inference. Its primary purpose is to facilitate the development of highly quantized models with an effective weight bit-width of 1.88 bits.

Loading preview...

MingZwhy/Qwen3-4B-W1.88-QAD: A Quantization-Aware-Distillation Checkpoint

This model, developed by MingZwhy, is a quantization-aware-distillation (QAD) checkpoint for the Qwen3-4B architecture, specifically configured for a W1.88 quantization scheme. It represents the initial state from which on-policy distillation (OPD) training commences.

Key Characteristics & Usage

  • Starting Point for Training: This checkpoint is not a ready-to-use model for direct inference. Instead, it's intended as a STUDENT_MODEL for the OPD stage, where the quantizer configuration is applied during training.
  • Latent Checkpoint: The tensors are in bf16 format and are not yet quantized. Quantization-aware training maintains high-precision master weights, with quantization applied within the forward pass. Loading this file directly yields an unquantized model.
  • Quantization Scheme: When used in the intended OPD process, it targets a mixed INT1.58 / INT4 weight quantization in blocks, resulting in an effective 1.88 bits for weights. Embedding and output heads are INT4, and activations are INT8.
  • Recommended Alternative: For a directly loadable and evaluable quantized model, users should refer to MingZwhy/Qwen3-4B-W1.88-QAOPD instead.

Intended Use Case

This model is specifically designed for researchers and developers engaged in quantization research and development, particularly those working on on-policy distillation to achieve highly efficient, low-bit models based on the Qwen3-4B architecture. It provides the foundational checkpoint for initiating such advanced quantization training workflows.