yueqis/swe_original-qwen-7b-25k-3epochs-5e-5
The yueqis/swe_original-qwen-7b-25k-3epochs-5e-5 model is a fine-tuned version of Qwen's Qwen2.5-7B-Instruct, featuring 7.6 billion parameters and a 32K context length. This model has been specifically adapted using the swe_original dataset, indicating a specialization in areas related to its training data. It is designed for tasks benefiting from its fine-tuning, offering enhanced performance in its targeted domain.
Loading preview...
Overview
This model, swe_original-qwen-7b-25k-3epochs-5e-5, is a fine-tuned iteration of the Qwen2.5-7B-Instruct base model developed by Qwen. It leverages a 7.6 billion parameter architecture with a substantial 32K token context length, making it suitable for processing longer sequences of text.
Key Characteristics
- Base Model: Qwen/Qwen2.5-7B-Instruct.
- Fine-tuning Dataset:
swe_originaldataset, suggesting a specialization in content related to this data. - Training Configuration: Trained for 3 epochs with a learning rate of 5e-05, using a total batch size of 128 across 8 GPUs. The optimizer used was
adamw_torchwith a cosine learning rate scheduler. - Performance: Achieved a loss of 0.0027 on the evaluation set, indicating effective learning during the fine-tuning process.
Intended Use Cases
This model is best suited for applications that align with the characteristics of the swe_original dataset it was fine-tuned on. Developers should consider its specific training for tasks requiring nuanced understanding or generation within that domain. Due to its instruction-tuned base and fine-tuning, it is likely to perform well in instruction-following scenarios relevant to its specialized data.