Sebastianpro88/Chichu-2.0-500M-Instruct
Sebastianpro88/Chichu-2.0-500M-Instruct is a 0.5 billion parameter instruction-tuned language model based on the Qwen2.5-0.5B-Instruct architecture. It was fine-tuned using LoRA on a multi-teacher distillation dataset covering math, code, reasoning, and general instructions. This model is designed for diverse instructional tasks, leveraging its compact size and 32,768 token context length for efficient deployment.
Loading preview...
Chichu 2.0 500M Instruct Overview
Chichu 2.0 500M Instruct is a compact yet capable language model developed by Sebastianpro88. It is built upon the Qwen2.5-0.5B-Instruct base model, featuring approximately 500 million parameters. The model was fine-tuned using LoRA (rank=16, alpha=32) with only 2.16 million LoRA adapters trained, making it efficient to adapt.
Key Capabilities & Training
- Base Architecture: Qwen2.5-0.5B-Instruct, providing a robust foundation.
- Parameter Efficiency: Utilizes LoRA fine-tuning, making it resource-efficient for deployment and further adaptation.
- Diverse Instruction Following: Trained on the
r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillationdataset, which is a multi-teacher distillation corpus. This dataset specifically targets a broad range of tasks including:- Mathematical problem-solving
- Code generation and understanding
- General reasoning abilities
- Following various instructions
- Extended Context Window: Supports a substantial context length of 32,768 tokens, allowing it to process and generate longer sequences of text.
Use Cases
This model is well-suited for applications requiring a small, efficient instruction-following model capable of handling tasks across math, code, and reasoning, especially where a large context window is beneficial.