miiikiik/Qwen3-8B-music-coa-sft
The miiikiik/Qwen3-8B-music-coa-sft model is an 8 billion parameter language model fine-tuned from a Qwen3-8B base model. It is specifically adapted using the music_coa_sft dataset, indicating a specialization in music-related content or tasks. This model is designed for applications requiring understanding or generation within the music domain, leveraging its 32768 token context length for processing longer sequences.
Loading preview...
Model Overview
This model, miiikiik/Qwen3-8B-music-coa-sft, is an 8 billion parameter language model built upon the Qwen3 architecture. It has been specifically fine-tuned from a base model located at /mnt/tidal-alsh01/usr/zhongmeizhi/meiwangyi/models/Qwen3-8B-onto-fixed-traj.
Key Specialization
The primary differentiator for this model is its fine-tuning on the music_coa_sft dataset. This suggests a specialization in tasks or content related to music, potentially including music theory, composition, analysis, or other music-centric applications.
Training Details
The model was trained with the following key hyperparameters:
- Learning Rate: 1e-05
- Batch Size: 2 (train), 8 (eval) with 4 gradient accumulation steps, leading to a total effective batch size of 64.
- Optimizer: ADAMW_TORCH
- LR Scheduler: Cosine with 0.1 warmup steps
- Epochs: 3.0
Intended Use
While specific intended uses and limitations require more information, its fine-tuning on a music-specific dataset indicates its suitability for applications within the music domain. Developers should consider its specialized training for tasks where music-related understanding or generation is crucial.