hahayhe/DReP-SFT-Qwen3-4B-Thinking-2507

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 3, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

hahayhe/DReP-SFT-Qwen3-4B-Thinking-2507 is a 4 billion parameter language model fine-tuned from the Qwen3-4B-Thinking-2507 base model. This model is specifically fine-tuned on the drep_a_half_sft dataset, indicating a specialization for tasks related to its training data. It is designed for general language understanding and generation within its 32768 token context window, with its specific strengths tied to the characteristics of the drep_a_half_sft dataset.

Loading preview...

Model Overview

hahayhe/DReP-SFT-Qwen3-4B-Thinking-2507 is a 4 billion parameter language model, fine-tuned from the Qwen3-4B-Thinking-2507 base model. This model has been specifically adapted through supervised fine-tuning (SFT) on the drep_a_half_sft dataset. The training process involved a learning rate of 1e-05, a total batch size of 128 across 4 GPUs, and a cosine learning rate scheduler with 0.1 warmup ratio over 4 epochs.

Key Training Details

  • Base Model: Qwen3-4B-Thinking-2507
  • Fine-tuning Dataset: drep_a_half_sft
  • Parameters: 4 billion
  • Context Length: 32768 tokens
  • Optimizer: AdamW with betas=(0.9, 0.999) and epsilon=1e-08
  • Epochs: 4
  • Frameworks: Transformers 4.57.1, Pytorch 2.6.0+cu124, Datasets 4.0.0, Tokenizers 0.22.2

Intended Use Cases

While specific intended uses and limitations require more detailed information from the model developer, this model is generally suitable for tasks aligned with the data it was fine-tuned on. Developers should evaluate its performance on their specific applications, particularly those that benefit from the characteristics of the drep_a_half_sft dataset.