huggingFacing/qwen2.5-7b-to-1.5b-liftkd-v8-bilingual100k-v2-continue-e2to4-step1500
The huggingFacing/qwen2.5-7b-to-1.5b-liftkd-v8-bilingual100k-v2-continue-e2to4-step1500 model is a 1.5 billion parameter instruction-tuned student model, distilled from the Qwen2.5-7B-Instruct model. It utilizes the LiftKD V8 method with GKD JSD and a normalized gap gate for efficient knowledge transfer. Trained on a 100K bilingual English/Chinese instruction mixture, including 18.75% mathematics, this model is optimized for instruction following in both languages. Its primary strength lies in its ability to perform well on instruction-based tasks despite its smaller size, making it suitable for resource-constrained applications requiring bilingual capabilities.
Loading preview...
Model Overview
This model, huggingFacing/qwen2.5-7b-to-1.5b-liftkd-v8-bilingual100k-v2-continue-e2to4-step1500, is a 1.5 billion parameter instruction-tuned student model. It was created through a knowledge distillation process from the larger Qwen2.5-7B-Instruct model.
Key Distillation and Training Details
- Distillation Method: Employs the LiftKD V8 technique, specifically using fully on-policy GKD JSD with a normalized gap gate for effective knowledge transfer.
- Training Data: Trained on a diverse 100,000-sample bilingual instruction mixture comprising both English and Chinese. Notably, 18.75% of this dataset is dedicated to mathematics-related instructions, enhancing its numerical reasoning capabilities.
- Training Configuration: Utilized BF16 full-parameter training with DeepSpeed ZeRO-2, a global batch size of 64, and a seed of 10.
- Sequence Limits: Configured for 384 prompt tokens, 512 total tokens, and 128 generated tokens during its training phase.
Use Cases and Differentiators
This model is particularly well-suited for applications requiring:
- Bilingual Instruction Following: Excels in understanding and responding to instructions in both English and Chinese.
- Resource-Efficient Deployment: Its 1.5B parameter count makes it a strong candidate for environments with limited computational resources, offering a distilled performance from a larger model.
- Mathematical Instruction Handling: The inclusion of a significant portion of mathematical data in its training mixture suggests improved performance on quantitative tasks.