Jani12067/qwen3-finetuned

TEXT GENERATIONConcurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 13, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Jani12067/qwen3-finetuned is a 0.8 billion parameter causal language model, fine-tuned from the Qwen/Qwen3-0.6B architecture. This model was fine-tuned on an unspecified dataset, achieving a validation loss of 2.0521. It is suitable for general language generation tasks where a smaller, fine-tuned model is preferred.

Loading preview...

Model Overview

Jani12067/qwen3-finetuned is a fine-tuned variant of the Qwen3-0.6B model, developed by Jani12067. This model has 0.8 billion parameters and a context length of 32768 tokens. It was trained for 3 epochs, achieving a final validation loss of 2.0521.

Key Characteristics

  • Base Model: Fine-tuned from Qwen/Qwen3-0.6B.
  • Parameter Count: 0.8 billion parameters, making it a relatively compact model.
  • Training Data: The specific dataset used for fine-tuning is not disclosed.
  • Performance Metric: Achieved a validation loss of 2.0521, indicating its performance on the evaluation set.

Training Details

The model was trained using the following hyperparameters:

  • Learning Rate: 2e-05
  • Batch Size: train_batch_size of 2, eval_batch_size of 8, with gradient_accumulation_steps of 8, resulting in a total_train_batch_size of 16.
  • Optimizer: ADAMW_TORCH_FUSED.
  • Epochs: Trained for 3 epochs.

Intended Uses

Given its fine-tuned nature and smaller parameter count, this model is suitable for applications requiring efficient language generation where the specific fine-tuning objective aligns with the use case. Further information on its intended uses and limitations would require details about the fine-tuning dataset and objectives.