etiennebamas/qwen3-step32k-1e5

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 26, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The etiennebamas/qwen3-step32k-1e5 model is an 8 billion parameter causal language model, fine-tuned from formalmathatepfl/qwen3-cpt, featuring a 32K context length. This model is specifically fine-tuned on an sft dataset, indicating an optimization for supervised fine-tuning tasks. Its primary differentiator lies in its fine-tuning process, making it suitable for applications requiring specialized instruction following or task completion based on the sft dataset.

Loading preview...

Model Overview

The etiennebamas/qwen3-step32k-1e5 is an 8 billion parameter language model, fine-tuned from the formalmathatepfl/qwen3-cpt base model. It supports a substantial context length of 32,768 tokens, enabling it to process and generate longer sequences of text. The model has undergone supervised fine-tuning (SFT) on a specific dataset, which suggests its capabilities are tailored towards tasks aligned with that training data.

Training Details

The model was trained for 1 epoch using a learning rate of 1e-05, with a total training batch size of 8 across 8 GPUs. It utilized the AdamW optimizer with a cosine learning rate scheduler and a warmup ratio of 0.05. The training environment included Transformers 4.57.3, Pytorch 2.9.0+cu128, Datasets 3.6.0, and Tokenizers 0.22.2.

Intended Uses

Given its fine-tuned nature, this model is likely best suited for:

  • Specific SFT-based tasks: Applications that align with the characteristics of the sft dataset it was trained on.
  • Long context understanding: Its 32K context window makes it suitable for tasks requiring comprehension or generation over extended text passages.

Further details on specific intended uses and limitations would require more information about the sft dataset used for fine-tuning.