Jani12067/qwen3-finetuned
Jani12067/qwen3-finetuned is a 0.8 billion parameter causal language model, fine-tuned from the Qwen/Qwen3-0.6B architecture. This model was fine-tuned on an unspecified dataset, achieving a validation loss of 2.0521. It is suitable for general language generation tasks where a smaller, fine-tuned model is preferred.
Loading preview...
Model Overview
Jani12067/qwen3-finetuned is a fine-tuned variant of the Qwen3-0.6B model, developed by Jani12067. This model has 0.8 billion parameters and a context length of 32768 tokens. It was trained for 3 epochs, achieving a final validation loss of 2.0521.
Key Characteristics
- Base Model: Fine-tuned from Qwen/Qwen3-0.6B.
- Parameter Count: 0.8 billion parameters, making it a relatively compact model.
- Training Data: The specific dataset used for fine-tuning is not disclosed.
- Performance Metric: Achieved a validation loss of 2.0521, indicating its performance on the evaluation set.
Training Details
The model was trained using the following hyperparameters:
- Learning Rate: 2e-05
- Batch Size:
train_batch_sizeof 2,eval_batch_sizeof 8, withgradient_accumulation_stepsof 8, resulting in atotal_train_batch_sizeof 16. - Optimizer: ADAMW_TORCH_FUSED.
- Epochs: Trained for 3 epochs.
Intended Uses
Given its fine-tuned nature and smaller parameter count, this model is suitable for applications requiring efficient language generation where the specific fine-tuning objective aligns with the use case. Further information on its intended uses and limitations would require details about the fine-tuning dataset and objectives.