arjunanand13/qwen3-finetuned
arjunanand13/qwen3-finetuned is a 0.8 billion parameter language model, fine-tuned from the Qwen/Qwen3-0.6B architecture. This model was trained for 1 epoch with a learning rate of 2e-05 and achieved a validation loss of 3.2289. Due to limited information on its training dataset and intended uses, its specific differentiators and primary applications are not explicitly defined.
Loading preview...
Model Overview
arjunanand13/qwen3-finetuned is a language model derived from the Qwen/Qwen3-0.6B base architecture. This version has undergone a fine-tuning process, resulting in a model with approximately 0.8 billion parameters. The specific dataset used for this fine-tuning is not detailed in the available information.
Training Details
The model was trained using the following hyperparameters:
- Learning Rate: 2e-05
- Batch Size: 16 (train), 8 (evaluation)
- Gradient Accumulation Steps: 16
- Total Train Batch Size: 256
- Optimizer: ADAMW_TORCH_FUSED
- LR Scheduler Type: Linear
- Epochs: 1
During training, the model achieved a validation loss of 3.2289. The training was conducted using Transformers 5.13.1, Pytorch 2.11.0+cu128, Datasets 4.0.0, and Tokenizers 0.22.2.
Current Limitations
Detailed information regarding the model's intended uses, specific capabilities, and limitations is not provided. Users should conduct further evaluation to determine its suitability for particular applications.