shreyamishra-05/qwen3-finetuned

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 8, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

shreyamishra-05/qwen3-finetuned is a 0.8 billion parameter language model, fine-tuned from the Qwen/Qwen3-0.6B architecture. This model was trained with a learning rate of 2e-05 over one epoch, achieving a validation loss of 3.2261. Its specific fine-tuning dataset and primary use cases are not detailed, suggesting a general-purpose application based on its base model's capabilities.

Loading preview...

Model Overview

shreyamishra-05/qwen3-finetuned is a fine-tuned language model based on the Qwen3-0.6B architecture, featuring approximately 0.8 billion parameters. This model was developed by shreyamishra-05 and represents an adaptation of the original Qwen3-0.6B base model.

Training Details

The model underwent a single epoch of fine-tuning with a learning rate of 2e-05. Key training hyperparameters included a train_batch_size of 16, eval_batch_size of 8, and a gradient_accumulation_steps of 16, resulting in a total_train_batch_size of 256. The optimizer used was ADAMW_TORCH_FUSED with standard beta values and an epsilon of 1e-08. During training, the model achieved a validation loss of 3.2261.

Current Limitations

Specific details regarding the fine-tuning dataset, intended uses, and potential limitations are not provided in the available documentation. Users should exercise caution and conduct further evaluation to determine its suitability for specific applications, as its unique differentiators and optimized use cases are not explicitly stated.