yueqis/swe_original-qwen-7b-30k-3epochs-5e-5
The yueqis/swe_original-qwen-7b-30k-3epochs-5e-5 model is a 7.6 billion parameter Qwen2.5-7B-Instruct variant, fine-tuned on the swe_original dataset. This model is optimized for tasks related to the specific data distribution of the swe_original dataset, demonstrating a final validation loss of 0.0029. It is suitable for applications requiring high accuracy on data similar to its training corpus, leveraging a 32K token context window.
Loading preview...
Model Overview
This model, yueqis/swe_original-qwen-7b-30k-3epochs-5e-5, is a fine-tuned version of the Qwen2.5-7B-Instruct base model, developed by Qwen. It features 7.6 billion parameters and supports a context length of 32,768 tokens. The model was specifically trained on the swe_original dataset over 3 epochs, achieving a final validation loss of 0.0029.
Key Training Details
- Base Model: Qwen/Qwen2.5-7B-Instruct
- Dataset:
swe_original - Learning Rate: 5e-05
- Optimizer: AdamW with betas=(0.9, 0.999)
- Batch Size: 1 (train), 1 (eval) with 16 gradient accumulation steps, resulting in a total train batch size of 128
- Epochs: 3
- Frameworks: Transformers 4.51.3, PyTorch 2.7.0+cu126, Datasets 3.5.0, Tokenizers 0.21.1
Intended Use Cases
This model is best suited for tasks that align closely with the data characteristics of the swe_original dataset. Its fine-tuning process aims to enhance performance and accuracy on specific data distributions, making it a strong candidate for applications where high fidelity to the training data's domain is crucial. Developers should consider its specialized training for tasks requiring nuanced understanding or generation within that domain.