rberberi/qwen3-finetuned

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Oct 4, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The rberberi/qwen3-finetuned model is a 0.8 billion parameter language model, fine-tuned by rberberi from the Qwen/Qwen3-0.6B base model. This model has a context length of 32768 tokens. While the specific fine-tuning dataset is unknown, its primary differentiator is its fine-tuned nature, suggesting optimization for specific, undisclosed tasks beyond the base model's general capabilities. It is intended for use cases that benefit from a compact, fine-tuned model with a substantial context window.

Loading preview...

Overview

The rberberi/qwen3-finetuned model is a specialized version of the Qwen3-0.6B base model, developed by rberberi. This model, with approximately 0.8 billion parameters, has been fine-tuned on an unspecified dataset, indicating a focus on particular tasks or domains not covered by the original base model. It supports a substantial context length of 32768 tokens, making it suitable for processing longer inputs.

Key Characteristics

  • Base Model: Fine-tuned from Qwen/Qwen3-0.6B.
  • Parameter Count: Approximately 0.8 billion parameters.
  • Context Length: Supports a large context window of 32768 tokens.
  • Fine-tuned: Optimized through fine-tuning, though the specific dataset and target tasks are not detailed in the provided information.

Training Details

The model was trained with a learning rate of 2e-05, a total batch size of 16 (achieved with train_batch_size: 2 and gradient_accumulation_steps: 8), and for 3 epochs. Mixed-precision training (Native AMP) was utilized. The training process resulted in a final validation loss of 1.2141.

Good for

  • Applications requiring a compact, fine-tuned language model.
  • Tasks benefiting from a large context window (32768 tokens).
  • Exploration of fine-tuned Qwen3 variants for specific, custom use cases.