etiennebamas/qwen3-lr-1-e-5-neat

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 27, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The etiennebamas/qwen3-lr-1-e-5-neat model is an 8 billion parameter language model fine-tuned from formalmathatepfl/qwen3-cpt, based on the Qwen3 architecture. This model was trained with a learning rate of 1e-05 and a context length of 32768 tokens. It is optimized for tasks related to its fine-tuning on the sft dataset, suggesting potential specialization in areas relevant to that data. Its large context window makes it suitable for processing extensive textual inputs.

Loading preview...

Model Overview

The etiennebamas/qwen3-lr-1-e-5-neat model is an 8 billion parameter language model, fine-tuned from the formalmathatepfl/qwen3-cpt base model. It leverages the Qwen3 architecture and was specifically trained on an 'sft' dataset, indicating a potential specialization derived from this fine-tuning process. The model supports a substantial context length of 32768 tokens, allowing it to process and generate responses based on very long inputs.

Key Training Details

Training was conducted with a learning rate of 1e-05, a batch size of 1 per device across 8 GPUs (totaling 8), and a cosine learning rate scheduler with a 0.05 warmup ratio over 1 epoch. The optimizer used was AdamW with standard betas and epsilon. This configuration suggests a focused fine-tuning approach aimed at adapting the base model to specific tasks or data characteristics.

Potential Use Cases

Given its fine-tuning on the 'sft' dataset and large context window, this model could be particularly useful for applications requiring:

  • Processing and understanding extensive documents or conversations.
  • Tasks aligned with the specific domain or nature of the 'sft' dataset.
  • Applications where a balance between model size and context handling is crucial.