AmanKhan2002/qwen3-finetuned

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 8, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

AmanKhan2002/qwen3-finetuned is a 0.8 billion parameter causal language model, fine-tuned by AmanKhan2002 from the Qwen3-0.6B architecture. This model was trained for 1 epoch with a learning rate of 2e-05 and achieved a validation loss of 3.2387. Its specific intended uses and primary differentiators are not detailed in the available information.

Loading preview...

Model Overview

This model, AmanKhan2002/qwen3-finetuned, is a fine-tuned variant of the Qwen3-0.6B base model developed by Qwen. It features approximately 0.8 billion parameters and was trained for a single epoch. The training process utilized a learning rate of 2e-05 with an AdamW optimizer, a batch size of 16, and a gradient accumulation of 16 steps, resulting in a total training batch size of 256. During evaluation, it achieved a validation loss of 3.2387.

Key Characteristics

  • Base Model: Qwen3-0.6B
  • Parameter Count: ~0.8 billion
  • Training Epochs: 1
  • Validation Loss: 3.2387
  • Frameworks: Built with Transformers 5.13.1, Pytorch 2.11.0+cu128, Datasets 4.0.0, and Tokenizers 0.22.2.

Intended Uses & Limitations

The specific intended uses, detailed capabilities, and limitations of this fine-tuned model are not explicitly provided in the available documentation. Further information would be needed to determine optimal applications or potential constraints.