huggingFacing/qwen2.5-7b-to-1.5b-liftkd-v8-bilingual100k-v2-continue-e2to4-final

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 16, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

This model is a 1.5 billion parameter Qwen2.5-Instruct student model, distilled from a larger Qwen2.5-7B-Instruct teacher model. Developed using the LiftKD V8 method, it was trained on a 100,000-sample bilingual English/Chinese instruction dataset, which includes a significant portion of mathematics data. It is optimized for efficient instruction-following in both English and Chinese, particularly for tasks involving mathematical reasoning.

Loading preview...

Model Overview

This model is a 1.5 billion parameter student model, derived from the more extensive Qwen2.5-7B-Instruct teacher model. It represents the cumulative epoch-4 checkpoint of a distillation process aimed at creating a smaller, efficient instruction-following model.

Key Distillation and Training Details

  • Distillation Method: The model was created using the LiftKD V8 method, specifically employing fully on-policy GKD JSD with a normalized gap gate. This advanced distillation technique helps transfer knowledge effectively from the larger teacher model.
  • Training Data: It was trained on a 100,000-sample bilingual instruction mixture comprising both English and Chinese data. Notably, 18.75% of this dataset is dedicated to mathematics, suggesting a focus on numerical and logical reasoning capabilities.
  • Training Configuration: The training utilized BF16 precision with full-parameter training and DeepSpeed ZeRO-2 for efficiency. A global batch size of 64 was used.
  • Sequence Limits: During training, the model processed prompts up to 384 tokens, with a total token limit of 512 and a generation limit of 128 tokens.

Potential Use Cases

  • Bilingual Instruction Following: Excels in understanding and responding to instructions in both English and Chinese.
  • Mathematical Reasoning: The inclusion of a substantial mathematics dataset suggests proficiency in handling mathematical queries and problems.
  • Resource-Constrained Environments: As a 1.5B parameter model, it offers a more efficient alternative to larger models while retaining significant capabilities due to its distillation from a 7B teacher.