Koki0511/qwen3-finetuned
Koki0511/qwen3-finetuned is a 0.8 billion parameter causal language model, fine-tuned from Qwen/Qwen3-0.6B. This model has a context length of 32768 tokens and was trained with a learning rate of 2e-05 over 3 epochs. It is a specialized version of the Qwen3 architecture, with its specific differentiators and primary use cases requiring further information from the developer.
Loading preview...
Model Overview
Koki0511/qwen3-finetuned is a 0.8 billion parameter language model, derived from the Qwen/Qwen3-0.6B base model. This fine-tuned version was developed using the transformers library and is licensed under Apache-2.0. The model's training involved a learning rate of 2e-05, a batch size of 2, and gradient accumulation steps of 8, totaling 3 epochs.
Training Details
During its training, the model achieved a validation loss of 2.0542. Key hyperparameters included an AdamW optimizer with default betas and epsilon, and a linear learning rate scheduler. The training process utilized Transformers 5.13.0, Pytorch 2.13.0+cu130, Datasets 5.0.0, and Tokenizers 0.22.2.
Current Status
As of its release, specific details regarding the dataset used for fine-tuning, the model's intended uses, limitations, and comprehensive evaluation results are not yet available in the provided documentation. Users should consult future updates for more information on its unique capabilities and optimal applications.