yueqis/swe_original-qwen-7b-30k
The yueqis/swe_original-qwen-7b-30k is a 7.6 billion parameter causal language model fine-tuned from Qwen/Qwen2.5-7B-Instruct. This model is specifically adapted using the swe_original dataset, indicating a specialization for tasks related to its training data. It is designed for applications requiring a Qwen2.5-7B-Instruct base model with further domain-specific fine-tuning.
Loading preview...
Model Overview
The yueqis/swe_original-qwen-7b-30k is a 7.6 billion parameter language model, fine-tuned from the robust Qwen/Qwen2.5-7B-Instruct base model. This fine-tuning process utilized the swe_original dataset, suggesting a specialization in areas relevant to this specific data.
Key Characteristics
- Base Model: Qwen2.5-7B-Instruct, known for its strong general language understanding and instruction-following capabilities.
- Parameter Count: 7.6 billion parameters, offering a balance between performance and computational efficiency.
- Context Length: Supports a substantial context window of 32,768 tokens, enabling processing of longer inputs and maintaining coherence over extended interactions.
- Fine-tuning Focus: The model has undergone specific fine-tuning on the
swe_originaldataset, which implies enhanced performance or tailored behavior for tasks aligned with this dataset's characteristics.
Training Details
The fine-tuning was conducted with a learning rate of 1e-05, a total batch size of 128 (achieved with gradient accumulation over 16 steps on 8 GPUs), and a cosine learning rate scheduler with a 0.05 warmup ratio over 1 epoch. The training achieved a loss of 0.1770 on the evaluation set.
Potential Use Cases
This model is suitable for applications that can leverage the strengths of the Qwen2.5-7B-Instruct architecture combined with the specific adaptations from the swe_original dataset. Developers looking for a model with a strong base and targeted fine-tuning for particular domain-specific tasks may find this model beneficial.