hahayhe/DReP-SFT-Qwen3-8B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 16, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The hahayhe/DReP-SFT-Qwen3-8B is an 8 billion parameter language model, fine-tuned from the Qwen3-8B architecture. This model is specifically trained on the drep_a_half_sft dataset, indicating a specialized fine-tuning process. It is designed for tasks aligned with its specific training data, offering focused performance rather than general-purpose capabilities. The model leverages a 32768 token context length, suitable for processing longer sequences of text.

Loading preview...

Model Overview

The hahayhe/DReP-SFT-Qwen3-8B is an 8 billion parameter language model, fine-tuned from the base Qwen3-8B architecture. This model has undergone specialized training on the drep_a_half_sft dataset, suggesting an optimization for tasks related to the characteristics of this specific data.

Training Details

The fine-tuning process utilized a learning rate of 1e-05, with a total training batch size of 128 across 4 devices. An AdamW optimizer with cosine learning rate scheduling and a warmup ratio of 0.1 was employed over 4 epochs. The model was trained using Transformers 4.57.1 and PyTorch 2.6.0+cu124.

Key Characteristics

  • Base Model: Qwen3-8B
  • Parameter Count: 8 billion
  • Context Length: 32768 tokens
  • Specialized Fine-tuning: Trained on the drep_a_half_sft dataset, indicating a focus on specific domain or task performance.

Intended Use

This model is best suited for applications that align with the data it was fine-tuned on. Developers should consider its specialized training for tasks where the drep_a_half_sft dataset's characteristics are relevant, rather than for broad, general-purpose language generation.