adityabanerjee13/qwen2.5-0.5b-cpt-mix-1to2
The adityabanerjee13/qwen2.5-0.5b-cpt-mix-1to2 model is a 0.5 billion parameter Qwen2.5-based causal language model fine-tuned by adityabanerjee13. It was trained with a specific 1:2 ratio of FineWeb to Indic character data, focusing on a mix of general web content and Indic language data. This model is optimized for tasks requiring understanding and generation across these mixed data distributions, leveraging its 32768 token context length.
Loading preview...
Model Overview
This model, adityabanerjee13/qwen2.5-0.5b-cpt-mix-1to2, is a 0.5 billion parameter variant of the Qwen2.5 architecture, fine-tuned by adityabanerjee13. It was trained using Axolotl with a unique data mix strategy, specifically a 1:2 character-level ratio of FineWeb web-crawl data to Indic language data.
Training Details
The training involved two primary datasets: adityabanerjee13/indic-cpt-mini-train and adityabanerjee13/fineweb-cpt-half, ensuring the precise 1:2 character ratio. The model utilizes a sequence length of 4096 tokens with sample packing enabled. Key hyperparameters include a learning rate of 2e-05, a total batch size of 32 (micro_batch_size 4, gradient_accumulation_steps 8), and 2 epochs of training. Evaluation was configured to log losses for Indic and FineWeb validation sets separately.
Key Characteristics
- Base Model: Qwen/Qwen2.5-0.5B.
- Data Mix: Fine-tuned on a 1:2 ratio of FineWeb to Indic character data.
- Context Length: Supports a sequence length of 4096 tokens.
- Training Framework: Built with Axolotl, enabling detailed configuration and multi-evaluation plugins.
Potential Use Cases
This model is suitable for applications requiring language understanding and generation that benefit from exposure to both general web content and Indic language data, particularly in scenarios where a smaller, efficient model is preferred.