huggingFacing/qwen2.5-7b-to-1.5b-liftkd-v8-bilingual100k-v2-continue-e2to4-step1500

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 16, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The huggingFacing/qwen2.5-7b-to-1.5b-liftkd-v8-bilingual100k-v2-continue-e2to4-step1500 model is a 1.5 billion parameter instruction-tuned student model, distilled from the Qwen2.5-7B-Instruct model. It utilizes the LiftKD V8 method with GKD JSD and a normalized gap gate for efficient knowledge transfer. Trained on a 100K bilingual English/Chinese instruction mixture, including 18.75% mathematics, this model is optimized for instruction following in both languages. Its primary strength lies in its ability to perform well on instruction-based tasks despite its smaller size, making it suitable for resource-constrained applications requiring bilingual capabilities.

Loading preview...

Model Overview

This model, huggingFacing/qwen2.5-7b-to-1.5b-liftkd-v8-bilingual100k-v2-continue-e2to4-step1500, is a 1.5 billion parameter instruction-tuned student model. It was created through a knowledge distillation process from the larger Qwen2.5-7B-Instruct model.

Key Distillation and Training Details

  • Distillation Method: Employs the LiftKD V8 technique, specifically using fully on-policy GKD JSD with a normalized gap gate for effective knowledge transfer.
  • Training Data: Trained on a diverse 100,000-sample bilingual instruction mixture comprising both English and Chinese. Notably, 18.75% of this dataset is dedicated to mathematics-related instructions, enhancing its numerical reasoning capabilities.
  • Training Configuration: Utilized BF16 full-parameter training with DeepSpeed ZeRO-2, a global batch size of 64, and a seed of 10.
  • Sequence Limits: Configured for 384 prompt tokens, 512 total tokens, and 128 generated tokens during its training phase.

Use Cases and Differentiators

This model is particularly well-suited for applications requiring:

  • Bilingual Instruction Following: Excels in understanding and responding to instructions in both English and Chinese.
  • Resource-Efficient Deployment: Its 1.5B parameter count makes it a strong candidate for environments with limited computational resources, offering a distilled performance from a larger model.
  • Mathematical Instruction Handling: The inclusion of a significant portion of mathematical data in its training mixture suggests improved performance on quantitative tasks.