adityabanerjee13/qwen2.5-0.5b-sft-IT
adityabanerjee13/qwen2.5-0.5b-sft-IT is a 0.5 billion parameter instruction-tuned causal language model developed by adityabanerjee13. It is a fine-tuned version of adityabanerjee13/qwen2.5-0.5b-cpt-mix-1to2, specifically trained on a mix of Indic and Tulu instruction-following datasets. This model is optimized for chat-based applications and instruction-following tasks, leveraging a 32768 token context length.
Loading preview...
Model Overview
adityabanerjee13/qwen2.5-0.5b-sft-IT is a 0.5 billion parameter instruction-tuned causal language model. It was developed by adityabanerjee13 and is a fine-tuned iteration of the adityabanerjee13/qwen2.5-0.5b-cpt-mix-1to2 base model. The fine-tuning process utilized a combination of the adityabanerjee13/indic-sft-mini-train and adityabanerjee13/tulu-sft-mini-train datasets, focusing on chat-template instruction following.
Key Training Details
- Base Model:
adityabanerjee13/qwen2.5-0.5b-cpt-mix-1to2 - Fine-tuning Datasets:
adityabanerjee13/indic-sft-mini-trainandadityabanerjee13/tulu-sft-mini-train - Training Framework: Built with Axolotl, version
0.19.0.dev0 - Sequence Length: 4096 tokens, with sample packing enabled.
- Optimization: AdamW_Torch_Fused optimizer, cosine learning rate scheduler with a learning rate of 2e-5.
- Context Length: The model supports a context length of 32768 tokens.
Intended Use Cases
This model is primarily designed for instruction-following and chat-based applications, particularly in contexts where the training data's characteristics (Indic and Tulu SFT) are relevant. Its compact size (0.5B parameters) makes it suitable for resource-constrained environments or applications requiring efficient inference.