Ayodeji711/qwen3-finetuned
Ayodeji711/qwen3-finetuned is a 0.8 billion parameter causal language model, fine-tuned from Qwen/Qwen3-0.6B. This model was trained with a 32768 token context length, achieving a validation loss of 2.0409. Its specific fine-tuning dataset and primary differentiators are not detailed, suggesting a general-purpose application or further exploration is needed to identify its specialized strengths.
Loading preview...
Model Overview
Ayodeji711/qwen3-finetuned is a 0.8 billion parameter language model derived from the Qwen/Qwen3-0.6B architecture. It has been fine-tuned over 3 epochs, achieving a final validation loss of 2.0409. The model was trained using a learning rate of 2e-05, a total batch size of 16 (with gradient accumulation steps of 8), and the AdamW_Torch_Fused optimizer.
Training Details
- Base Model: Qwen/Qwen3-0.6B
- Parameter Count: 0.8 billion
- Context Length: 32768 tokens
- Optimizer: AdamW_Torch_Fused with betas=(0.9, 0.999) and epsilon=1e-08
- Learning Rate: 2e-05
- Epochs: 3
- Final Validation Loss: 2.0409
Limitations
The specific dataset used for fine-tuning is not disclosed, and detailed information regarding the model's intended uses, limitations, and training data is currently unavailable. Users should exercise caution and conduct further evaluation to determine its suitability for specific applications.