yueqis/swe_original-qwen-7b-25k-3epochs-5e-5

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 15, 2025License:otherArchitecture:Transformer Featherless Exclusive Cold

The yueqis/swe_original-qwen-7b-25k-3epochs-5e-5 model is a fine-tuned version of Qwen's Qwen2.5-7B-Instruct, featuring 7.6 billion parameters and a 32K context length. This model has been specifically adapted using the swe_original dataset, indicating a specialization in areas related to its training data. It is designed for tasks benefiting from its fine-tuning, offering enhanced performance in its targeted domain.

Loading preview...

Overview

This model, swe_original-qwen-7b-25k-3epochs-5e-5, is a fine-tuned iteration of the Qwen2.5-7B-Instruct base model developed by Qwen. It leverages a 7.6 billion parameter architecture with a substantial 32K token context length, making it suitable for processing longer sequences of text.

Key Characteristics

  • Base Model: Qwen/Qwen2.5-7B-Instruct.
  • Fine-tuning Dataset: swe_original dataset, suggesting a specialization in content related to this data.
  • Training Configuration: Trained for 3 epochs with a learning rate of 5e-05, using a total batch size of 128 across 8 GPUs. The optimizer used was adamw_torch with a cosine learning rate scheduler.
  • Performance: Achieved a loss of 0.0027 on the evaluation set, indicating effective learning during the fine-tuning process.

Intended Use Cases

This model is best suited for applications that align with the characteristics of the swe_original dataset it was fine-tuned on. Developers should consider its specific training for tasks requiring nuanced understanding or generation within that domain. Due to its instruction-tuned base and fine-tuning, it is likely to perform well in instruction-following scenarios relevant to its specialized data.