imessam/qwen3-finetuned

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Oct 6, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

imessam/qwen3-finetuned is a 0.8 billion parameter language model fine-tuned from Qwen/Qwen3-0.6B. This model has a context length of 32768 tokens. Specific details regarding its training dataset, primary differentiators, and intended use cases are not provided in the available documentation. It is a base model that has undergone further training, with a reported validation loss of 2.0762.

Loading preview...

Overview

imessam/qwen3-finetuned is a language model derived from Qwen/Qwen3-0.6B, a model developed by Qwen. This version has been further fine-tuned, though the specific dataset used for this process is currently unknown. The model has 0.8 billion parameters and supports a context length of 32768 tokens.

Training Details

The fine-tuning process involved specific hyperparameters:

  • Learning Rate: 2e-05
  • Batch Size: 2 (train), 8 (eval)
  • Gradient Accumulation Steps: 8, leading to a total train batch size of 16
  • Optimizer: ADAMW_TORCH_FUSED
  • LR Scheduler: Linear
  • Epochs: 3

During training, the model achieved a final validation loss of 2.0762. The training was conducted using Transformers 5.18.0, Pytorch 2.14.1+cu130, Datasets 5.0.1, and Tokenizers 0.23.2.

Limitations

Detailed information regarding the model's intended uses, specific capabilities, and limitations is not available in the current documentation. Users should be aware that the dataset used for fine-tuning is unspecified, which may impact its performance on particular tasks.