yueqis/swe_original-qwen-7b-30k

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 13, 2025License:otherArchitecture:Transformer Featherless Exclusive Cold

The yueqis/swe_original-qwen-7b-30k is a 7.6 billion parameter causal language model fine-tuned from Qwen/Qwen2.5-7B-Instruct. This model is specifically adapted using the swe_original dataset, indicating a specialization for tasks related to its training data. It is designed for applications requiring a Qwen2.5-7B-Instruct base model with further domain-specific fine-tuning.

Loading preview...

Model Overview

The yueqis/swe_original-qwen-7b-30k is a 7.6 billion parameter language model, fine-tuned from the robust Qwen/Qwen2.5-7B-Instruct base model. This fine-tuning process utilized the swe_original dataset, suggesting a specialization in areas relevant to this specific data.

Key Characteristics

  • Base Model: Qwen2.5-7B-Instruct, known for its strong general language understanding and instruction-following capabilities.
  • Parameter Count: 7.6 billion parameters, offering a balance between performance and computational efficiency.
  • Context Length: Supports a substantial context window of 32,768 tokens, enabling processing of longer inputs and maintaining coherence over extended interactions.
  • Fine-tuning Focus: The model has undergone specific fine-tuning on the swe_original dataset, which implies enhanced performance or tailored behavior for tasks aligned with this dataset's characteristics.

Training Details

The fine-tuning was conducted with a learning rate of 1e-05, a total batch size of 128 (achieved with gradient accumulation over 16 steps on 8 GPUs), and a cosine learning rate scheduler with a 0.05 warmup ratio over 1 epoch. The training achieved a loss of 0.1770 on the evaluation set.

Potential Use Cases

This model is suitable for applications that can leverage the strengths of the Qwen2.5-7B-Instruct architecture combined with the specific adaptations from the swe_original dataset. Developers looking for a model with a strong base and targeted fine-tuning for particular domain-specific tasks may find this model beneficial.