etiennebamas/qwen3-sft-feedback-small-dataset

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 1, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

etiennebamas/qwen3-sft-feedback-small-dataset is an 8 billion parameter language model, fine-tuned from formalmathatepfl/qwen3-cpt. This model is specifically adapted through supervised fine-tuning (SFT) on a feedback dataset, aiming to enhance its response generation based on human feedback. With a context length of 32768 tokens, it is designed for tasks requiring improved conversational quality and alignment.

Loading preview...

Model Overview

etiennebamas/qwen3-sft-feedback-small-dataset is an 8 billion parameter language model, fine-tuned from the formalmathatepfl/qwen3-cpt base model. This version has undergone supervised fine-tuning (SFT) using a feedback dataset, indicating an optimization for generating responses that are more aligned with human preferences or specific feedback criteria.

Key Characteristics

  • Base Model: Fine-tuned from formalmathatepfl/qwen3-cpt.
  • Parameter Count: 8 billion parameters.
  • Context Length: Supports a substantial context window of 32768 tokens.
  • Training Method: Utilizes supervised fine-tuning (SFT) on a feedback dataset, suggesting an emphasis on improving response quality and alignment.

Training Details

The model was trained with a learning rate of 2e-05, a batch size of 1 per device across 8 GPUs (totaling 8), and for 1 epoch. The optimizer used was AdamW with cosine learning rate scheduling and a warmup ratio of 0.05. The training leveraged Transformers 4.57.3 and PyTorch 2.9.0+cu128.

Intended Use Cases

While specific intended uses are not detailed, the SFT on a feedback dataset implies suitability for applications where response quality, alignment with user intent, and conversational flow are critical. This could include chatbots, dialogue systems, or tasks requiring nuanced language generation based on iterative feedback.