abubakarilyas624/qwen3-finetuned

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 28, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The abubakarilyas624/qwen3-finetuned model is a fine-tuned version of the Qwen3-0.6B architecture, featuring approximately 0.8 billion parameters and a 32768-token context length. This model has been fine-tuned over 3 epochs, achieving a final validation loss of 2.0463. While the specific dataset and intended uses are not detailed, it represents a specialized adaptation of the Qwen3 base model.

Loading preview...

Model Overview

This model, abubakarilyas624/qwen3-finetuned, is a specialized adaptation built upon the Qwen3-0.6B architecture. It features approximately 0.8 billion parameters and supports a substantial 32768-token context length, indicating its potential for handling long sequences of text.

Training Details

The model underwent a fine-tuning process over 3 epochs using specific hyperparameters:

  • Learning Rate: 2e-05
  • Batch Size: 2 (train), 8 (eval) with 8 gradient accumulation steps, resulting in a total effective batch size of 16.
  • Optimizer: AdamW with betas=(0.9, 0.999) and epsilon=1e-08.
  • Scheduler: Linear learning rate scheduler.

Performance

During training, the model achieved a final validation loss of 2.0463, with the training loss progressively decreasing from 2.3766 in the first epoch to 1.9663 in the third.

Limitations

Crucially, the README indicates that more information is needed regarding the specific training and evaluation data, as well as its intended uses and limitations. Users should be aware that without this information, the model's specific strengths and appropriate applications are not clearly defined.