etiennebamas/qwen3-sft-classic-small-data

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 1, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

etiennebamas/qwen3-sft-classic-small-data is an 8 billion parameter language model, fine-tuned from formalmathatepfl/qwen3-cpt on a specific SFT dataset. This model is designed for general language tasks, leveraging its Qwen3 base architecture and a 32K context length. Its fine-tuning process aims to enhance performance on instruction-following and conversational applications.

Loading preview...

Model Overview

This model, etiennebamas/qwen3-sft-classic-small-data, is an 8 billion parameter language model derived from formalmathatepfl/qwen3-cpt. It has undergone supervised fine-tuning (SFT) using a dedicated dataset, aiming to improve its capabilities in instruction-following and general conversational tasks. The base model's architecture, Qwen3, provides a robust foundation for diverse language generation and understanding.

Key Characteristics

  • Base Model: Fine-tuned from formalmathatepfl/qwen3-cpt.
  • Parameter Count: 8 billion parameters, offering a balance between performance and computational efficiency.
  • Context Length: Supports a substantial context window of 32,768 tokens, enabling processing of longer inputs and generating coherent, extended responses.
  • Training: Fine-tuned with a learning rate of 2e-05, using an AdamW optimizer and a cosine learning rate scheduler over 1 epoch.

Intended Use Cases

This model is suitable for applications requiring a capable language model with a focus on instruction adherence and generating human-like text. Its fine-tuning on an SFT dataset suggests improved performance in:

  • Instruction Following: Responding accurately to given prompts and instructions.
  • Conversational AI: Engaging in more natural and coherent dialogues.
  • Text Generation: Creating various forms of text content based on input.

Further details on specific intended uses and limitations would require more information from the original model card.