huggingFacing/qwen2.5-7b-to-1.5b-liftkd-v8-bilingual100k-v2-continue-e2to4-step3000
The huggingFacing/qwen2.5-7b-to-1.5b-liftkd-v8-bilingual100k-v2-continue-e2to4-step3000 is a 1.5 billion parameter Qwen2.5-Instruct student model, distilled from the larger Qwen2.5-7B-Instruct. This model was trained using the LiftKD V8 method on a 100K bilingual English/Chinese instruction mixture, including a significant portion of mathematics data. It is optimized for efficient instruction-following in both English and Chinese, particularly for tasks within its sequence limits of 384 prompt tokens and 128 generated tokens.
Loading preview...
Model Overview
This model, huggingFacing/qwen2.5-7b-to-1.5b-liftkd-v8-bilingual100k-v2-continue-e2to4-step3000, is a 1.5 billion parameter student model derived from the Qwen2.5-7B-Instruct architecture. It represents the cumulative epoch-3 checkpoint of a distillation process aimed at creating a smaller, efficient instruction-following model.
Key Characteristics
- Distillation Method: Utilizes the LiftKD V8 technique, specifically employing fully on-policy GKD JSD with a normalized gap gate for knowledge transfer.
- Training Data: Trained on a 100,000-entry bilingual instruction mixture comprising both English and Chinese data. Notably, 18.75% of this dataset is dedicated to mathematics-related instructions.
- Sequence Limits: Designed for specific sequence lengths, supporting up to 384 prompt tokens, a total of 512 tokens, and generating up to 128 tokens.
- Precision: Trained using BF16 full-parameter training in conjunction with DeepSpeed ZeRO-2.
Use Cases
This model is particularly well-suited for applications requiring a compact, bilingual (English/Chinese) instruction-following LLM, especially where computational resources are a consideration. Its training on a mathematics-inclusive dataset suggests potential for tasks involving numerical reasoning within its token generation limits. It's ideal for scenarios where the efficiency of a 1.5B parameter model is preferred over larger alternatives, while still retaining instruction-following capabilities distilled from a 7B model.