Lakshay747/qwen3-finetuned

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 10, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Lakshay747/qwen3-finetuned is a 0.8 billion parameter causal language model, fine-tuned from the Qwen/Qwen3-0.6B architecture. This model was trained for 1 epoch with a learning rate of 2e-05 and achieved a validation loss of 3.1131. Due to limited information on its training dataset and specific optimizations, its primary differentiators and intended use cases are not explicitly defined beyond its base Qwen3 capabilities.

Loading preview...

Model Overview

Lakshay747/qwen3-finetuned is a language model derived from the Qwen/Qwen3-0.6B base architecture. This version has undergone a fine-tuning process, resulting in a 0.8 billion parameter model. The specific dataset used for fine-tuning is not detailed in the available information.

Training Details

The model was trained using the following hyperparameters:

  • Learning Rate: 2e-05
  • Batch Size: 8 (train and eval)
  • Gradient Accumulation Steps: 8, leading to a total train batch size of 64
  • Optimizer: ADAMW_TORCH_FUSED
  • LR Scheduler Type: Linear
  • Epochs: 1

During training, the model achieved a final validation loss of 3.1131. The training was conducted using Transformers 5.13.1, Pytorch 2.11.0+cu128, Datasets 4.0.0, and Tokenizers 0.22.2.

Current Limitations

Detailed information regarding the model's specific intended uses, limitations, and the characteristics of its training and evaluation data is not provided. Users should consider this when evaluating its suitability for particular applications.