adityabanerjee13/qwen2.5-0.5b-sft-IT

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 30, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

adityabanerjee13/qwen2.5-0.5b-sft-IT is a 0.5 billion parameter instruction-tuned causal language model developed by adityabanerjee13. It is a fine-tuned version of adityabanerjee13/qwen2.5-0.5b-cpt-mix-1to2, specifically trained on a mix of Indic and Tulu instruction-following datasets. This model is optimized for chat-based applications and instruction-following tasks, leveraging a 32768 token context length.

Loading preview...

Model Overview

adityabanerjee13/qwen2.5-0.5b-sft-IT is a 0.5 billion parameter instruction-tuned causal language model. It was developed by adityabanerjee13 and is a fine-tuned iteration of the adityabanerjee13/qwen2.5-0.5b-cpt-mix-1to2 base model. The fine-tuning process utilized a combination of the adityabanerjee13/indic-sft-mini-train and adityabanerjee13/tulu-sft-mini-train datasets, focusing on chat-template instruction following.

Key Training Details

  • Base Model: adityabanerjee13/qwen2.5-0.5b-cpt-mix-1to2
  • Fine-tuning Datasets: adityabanerjee13/indic-sft-mini-train and adityabanerjee13/tulu-sft-mini-train
  • Training Framework: Built with Axolotl, version 0.19.0.dev0
  • Sequence Length: 4096 tokens, with sample packing enabled.
  • Optimization: AdamW_Torch_Fused optimizer, cosine learning rate scheduler with a learning rate of 2e-5.
  • Context Length: The model supports a context length of 32768 tokens.

Intended Use Cases

This model is primarily designed for instruction-following and chat-based applications, particularly in contexts where the training data's characteristics (Indic and Tulu SFT) are relevant. Its compact size (0.5B parameters) makes it suitable for resource-constrained environments or applications requiring efficient inference.