huggingFacing/qwen2.5-7b-to-1.5b-liftkd-v8-bilingual100k-v2-continue-e2to4-step3000

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 16, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The huggingFacing/qwen2.5-7b-to-1.5b-liftkd-v8-bilingual100k-v2-continue-e2to4-step3000 is a 1.5 billion parameter Qwen2.5-Instruct student model, distilled from the larger Qwen2.5-7B-Instruct. This model was trained using the LiftKD V8 method on a 100K bilingual English/Chinese instruction mixture, including a significant portion of mathematics data. It is optimized for efficient instruction-following in both English and Chinese, particularly for tasks within its sequence limits of 384 prompt tokens and 128 generated tokens.

Loading preview...

Model Overview

This model, huggingFacing/qwen2.5-7b-to-1.5b-liftkd-v8-bilingual100k-v2-continue-e2to4-step3000, is a 1.5 billion parameter student model derived from the Qwen2.5-7B-Instruct architecture. It represents the cumulative epoch-3 checkpoint of a distillation process aimed at creating a smaller, efficient instruction-following model.

Key Characteristics

  • Distillation Method: Utilizes the LiftKD V8 technique, specifically employing fully on-policy GKD JSD with a normalized gap gate for knowledge transfer.
  • Training Data: Trained on a 100,000-entry bilingual instruction mixture comprising both English and Chinese data. Notably, 18.75% of this dataset is dedicated to mathematics-related instructions.
  • Sequence Limits: Designed for specific sequence lengths, supporting up to 384 prompt tokens, a total of 512 tokens, and generating up to 128 tokens.
  • Precision: Trained using BF16 full-parameter training in conjunction with DeepSpeed ZeRO-2.

Use Cases

This model is particularly well-suited for applications requiring a compact, bilingual (English/Chinese) instruction-following LLM, especially where computational resources are a consideration. Its training on a mathematics-inclusive dataset suggests potential for tasks involving numerical reasoning within its token generation limits. It's ideal for scenarios where the efficiency of a 1.5B parameter model is preferred over larger alternatives, while still retaining instruction-following capabilities distilled from a 7B model.