yueqis/full_sft_non_web-qwen-7b-25k-3epochs-5e-5

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 15, 2025License:otherArchitecture:Transformer Featherless Exclusive Cold

The yueqis/full_sft_non_web-qwen-7b-25k-3epochs-5e-5 model is a 7.6 billion parameter language model, fine-tuned from Qwen/Qwen2.5-7B-Instruct. It was trained on the full_sft_non_web dataset over 3 epochs with a learning rate of 5e-05. This model is optimized for tasks related to its specific fine-tuning dataset, achieving a loss of 0.2291 on the evaluation set.

Loading preview...

Overview

This model, full_sft_non_web-qwen-7b-25k-3epochs-5e-5, is a fine-tuned variant of the Qwen/Qwen2.5-7B-Instruct base model. It leverages a 7.6 billion parameter architecture and was specifically trained on the full_sft_non_web dataset. The training process involved 3 epochs with a learning rate of 5e-05, utilizing a multi-GPU setup with 8 devices and a total batch size of 128.

Key Training Details

  • Base Model: Qwen/Qwen2.5-7B-Instruct
  • Dataset: full_sft_non_web
  • Learning Rate: 5e-05
  • Epochs: 3
  • Optimizer: AdamW with cosine learning rate scheduler
  • Evaluation Loss: 0.2291

Intended Use Cases

Given its fine-tuning on the full_sft_non_web dataset, this model is best suited for applications and tasks that align with the characteristics and content of that specific dataset. Developers should consider its specialized training for use cases where the full_sft_non_web data is relevant, as its performance is evaluated based on this specific context.