eggbiscuit/Qwen2.5-0.5B-Distilled-from-1.5B-AlpacaChinese-v1

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:May 4, 2025License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

The eggbiscuit/Qwen2.5-0.5B-Distilled-from-1.5B-AlpacaChinese-v1 is a 0.5 billion parameter Chinese-language model based on the Qwen2.5 architecture, distilled from Qwen2.5-1.5B-Instruct. It utilizes a joint distillation framework with soft and hard labels to balance performance and efficiency. This model is optimized for lightweight deployment on edge devices, offering enhanced Chinese proficiency across tasks like Q&A, reasoning, and writing, with a context length of 32768 tokens.

Loading preview...

Model Overview

The eggbiscuit/Qwen2.5-0.5B-Distilled-from-1.5B-AlpacaChinese-v1 is a compact 0.5 billion parameter Chinese language model. It is a distilled version of the larger Qwen2.5-1.5B-Instruct model, trained by eggbiscuit using the Alpaca-Chinese-5K-Distill-Qwen dataset. This distillation process employs a unique joint optimization strategy combining KL Divergence (soft labels) and Cross-Entropy Loss (hard labels) to efficiently transfer capabilities from the teacher model while maintaining a small footprint.

Key Capabilities & Features

  • Lightweight and Efficient: Designed for resource-constrained environments, making it suitable for local and edge device deployment.
  • Enhanced Chinese Proficiency: Excels in various Chinese-language tasks, including question answering, reasoning, and text generation.
  • Joint Distillation: Leverages a sophisticated distillation method to preserve the performance characteristics of its larger teacher model.
  • Quantization-Ready: Fully supports q8_0 quantization, demonstrated with smooth performance on devices like iPhone 15 Pro Max and MacBook Air M1.

Performance Highlights

On the Chinese C-Eval benchmark, this model achieves an average score of 51.0, outperforming LLaMA2-65B (43.7) and closely matching Qwen-1.8B (50.3), despite being significantly smaller. Its efficiency is further highlighted by deployment tests, showing token speeds of 175.32 tok/s (prompt) / 73.29 tok/s (generation) on an iPhone 15 Pro Max.

Ideal Use Cases

This model is particularly well-suited for applications requiring a highly efficient, Chinese-proficient language model that can run directly on consumer-grade hardware or edge devices, where computational resources are limited but strong language understanding and generation capabilities are needed.