imessam/qwen3-finetuned
imessam/qwen3-finetuned is a 0.8 billion parameter language model fine-tuned from Qwen/Qwen3-0.6B. This model has a context length of 32768 tokens. Specific details regarding its training dataset, primary differentiators, and intended use cases are not provided in the available documentation. It is a base model that has undergone further training, with a reported validation loss of 2.0762.
Loading preview...
Overview
imessam/qwen3-finetuned is a language model derived from Qwen/Qwen3-0.6B, a model developed by Qwen. This version has been further fine-tuned, though the specific dataset used for this process is currently unknown. The model has 0.8 billion parameters and supports a context length of 32768 tokens.
Training Details
The fine-tuning process involved specific hyperparameters:
- Learning Rate: 2e-05
- Batch Size: 2 (train), 8 (eval)
- Gradient Accumulation Steps: 8, leading to a total train batch size of 16
- Optimizer: ADAMW_TORCH_FUSED
- LR Scheduler: Linear
- Epochs: 3
During training, the model achieved a final validation loss of 2.0762. The training was conducted using Transformers 5.18.0, Pytorch 2.14.1+cu130, Datasets 5.0.1, and Tokenizers 0.23.2.
Limitations
Detailed information regarding the model's intended uses, specific capabilities, and limitations is not available in the current documentation. Users should be aware that the dataset used for fine-tuning is unspecified, which may impact its performance on particular tasks.