huggingFacing/qwen2.5-7b-to-1.5b-liftkd-v12-top-only-paper100k-step1500-seed10

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 17, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

This model is a 1.5 billion parameter Qwen2.5-based language model, derived from a 7B teacher model through knowledge distillation. It utilizes a specific 'gap gate with step-level online top-two-layer influence weights only' variant for efficient compression. Optimized for maintaining performance in a smaller footprint, it is suitable for applications requiring a compact yet capable LLM.

Loading preview...

Model Overview

This model, huggingFacing/qwen2.5-7b-to-1.5b-liftkd-v12-top-only-paper100k-step1500-seed10, is a 1.5 billion parameter language model resulting from a knowledge distillation process. It was created by distilling knowledge from a larger Qwen/Qwen2.5-7B-Instruct teacher model into a Qwen/Qwen2.5-1.5B-Instruct student initialization.

Key Distillation Details

  • Variant: Employs a 'gap gate with step-level online top-two-layer influence weights only' method for knowledge transfer.
  • Training Data: Utilized lift_paper_en_natural_v1/100k, comprising 96,000 training examples and a 2,000-example controller split.
  • Objective: Trained using fully on-policy GKD (Generative Knowledge Distillation) over 1,500 optimizer steps.
  • Optimizer: AdamW with a cosine learning rate schedule from 1e-5 to 1e-7 and 1e-2 weight decay.

Use Cases

This model is designed for scenarios where a smaller, more efficient language model is required, while aiming to retain a significant portion of the capabilities of its larger 7B teacher. Its compact size makes it suitable for deployment in environments with limited computational resources or for applications demanding faster inference times.