baban/QwenTranslate_English_Bengali_100K_SFT
The baban/QwenTranslate_English_Bengali_100K_SFT model is a fine-tuned version of the Qwen/Qwen2.5-3B-Instruct architecture, developed by baban. This model specializes in English-to-Bengali translation tasks, leveraging supervised fine-tuning on the MT_En_Bengali dataset. It is optimized for achieving low loss in translation, making it suitable for applications requiring accurate language conversion between English and Bengali.
Loading preview...
Model Overview
This model, baban/QwenTranslate_English_Bengali_100K_SFT, is a specialized fine-tuned variant of the Qwen/Qwen2.5-3B-Instruct base model. Its primary function is English-to-Bengali translation, having undergone supervised fine-tuning on the MT_En_Bengali dataset.
Key Capabilities
- English-Bengali Translation: Specifically trained for translating text from English to Bengali.
- Fine-tuned Performance: Achieved a validation loss of 0.5930 on the evaluation set, indicating its proficiency in the target translation task.
Training Details
The model was trained using the following key hyperparameters:
- Learning Rate: 5e-05
- Optimizer: AdamW with betas=(0.9, 0.999) and epsilon=1e-08
- Batch Size: A total training batch size of 1024 (with gradient accumulation steps of 16 across 8 GPUs).
- Epochs: Trained for 3.0 epochs.
Intended Uses
This model is particularly well-suited for applications requiring dedicated English-to-Bengali translation capabilities, such as:
- Machine translation systems for English and Bengali.
- Content localization efforts targeting Bengali speakers.
- Research and development in low-resource language translation, specifically for the Bengali language pair.