huggingFacing/qwen2.5-7b-to-1.5b-liftkd-v8-bilingual100k-v2-continue-e2to4-final
TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 16, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold
This model is a 1.5 billion parameter Qwen2.5-Instruct student model, distilled from a larger Qwen2.5-7B-Instruct teacher model. Developed using the LiftKD V8 method, it was trained on a 100,000-sample bilingual English/Chinese instruction dataset, which includes a significant portion of mathematics data. It is optimized for efficient instruction-following in both English and Chinese, particularly for tasks involving mathematical reasoning.
Loading preview...
Model Overview
This model is a 1.5 billion parameter student model, derived from the more extensive Qwen2.5-7B-Instruct teacher model. It represents the cumulative epoch-4 checkpoint of a distillation process aimed at creating a smaller, efficient instruction-following model.
Key Distillation and Training Details
- Distillation Method: The model was created using the LiftKD V8 method, specifically employing fully on-policy GKD JSD with a normalized gap gate. This advanced distillation technique helps transfer knowledge effectively from the larger teacher model.
- Training Data: It was trained on a 100,000-sample bilingual instruction mixture comprising both English and Chinese data. Notably, 18.75% of this dataset is dedicated to mathematics, suggesting a focus on numerical and logical reasoning capabilities.
- Training Configuration: The training utilized BF16 precision with full-parameter training and DeepSpeed ZeRO-2 for efficiency. A global batch size of 64 was used.
- Sequence Limits: During training, the model processed prompts up to 384 tokens, with a total token limit of 512 and a generation limit of 128 tokens.
Potential Use Cases
- Bilingual Instruction Following: Excels in understanding and responding to instructions in both English and Chinese.
- Mathematical Reasoning: The inclusion of a substantial mathematics dataset suggests proficiency in handling mathematical queries and problems.
- Resource-Constrained Environments: As a 1.5B parameter model, it offers a more efficient alternative to larger models while retaining significant capabilities due to its distillation from a 7B teacher.