OP12138/qwen3-4b-star1
OP12138/qwen3-4b-star1 is a 4 billion parameter causal language model, fine-tuned from the Qwen3-4B architecture. This model is specifically adapted using the 'star1' dataset, indicating a specialized application or domain. It is designed for tasks aligned with its fine-tuning data, offering focused performance within that context.
Loading preview...
Model Overview
OP12138/qwen3-4b-star1 is a 4 billion parameter language model, fine-tuned from the base Qwen/Qwen3-4B architecture. This model has undergone specific adaptation using the 'star1' dataset, suggesting a specialization for particular tasks or data distributions.
Training Details
The fine-tuning process involved the following key hyperparameters:
- Learning Rate: 1e-05
- Batch Size: A total effective batch size of 16 (train_batch_size: 2, gradient_accumulation_steps: 8)
- Optimizer: Paged AdamW 8-bit with default betas and epsilon
- Scheduler: Cosine learning rate scheduler with a 0.05 warmup ratio
- Epochs: Trained for 5.0 epochs
This configuration indicates a focused fine-tuning approach to adapt the base Qwen3-4B model to the characteristics of the 'star1' dataset. Developers should consider the nature of the 'star1' dataset when evaluating its suitability for their specific use cases.