SuperXyrex/qwen3-finetuned
SuperXyrex/qwen3-finetuned is a 0.8 billion parameter language model, fine-tuned from Qwen/Qwen3-0.6B. This model was trained on an unspecified dataset, achieving a validation loss of 2.0564. It is a general-purpose language model, with specific applications and limitations yet to be detailed by the developer.
Loading preview...
Model Overview
SuperXyrex/qwen3-finetuned is a language model based on the Qwen3-0.6B architecture, developed by SuperXyrex. This version has been fine-tuned from the original Qwen/Qwen3-0.6B model, featuring 0.8 billion parameters and a context length of 32768 tokens.
Training Details
The model was trained over 3 epochs using a learning rate of 2e-05 and a total batch size of 16 (train_batch_size: 2, gradient_accumulation_steps: 8). The optimizer used was ADAMW_TORCH_FUSED. During training, the validation loss decreased from 2.3622 in the first epoch to 2.0564 by the third epoch.
Key Characteristics
- Base Model: Fine-tuned from Qwen/Qwen3-0.6B.
- Parameter Count: 0.8 billion parameters.
- Context Window: Supports a context length of 32768 tokens.
- Training Objective: Achieved a final validation loss of 2.0564.
Intended Use Cases
Specific intended uses and limitations for this fine-tuned model are not yet detailed in the provided information. Developers should consider its base architecture and training loss when evaluating its suitability for general language tasks.