miiikiik/Qwen3-8B-music-coa-sft

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 15, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The miiikiik/Qwen3-8B-music-coa-sft model is an 8 billion parameter language model fine-tuned from a Qwen3-8B base model. It is specifically adapted using the music_coa_sft dataset, indicating a specialization in music-related content or tasks. This model is designed for applications requiring understanding or generation within the music domain, leveraging its 32768 token context length for processing longer sequences.

Loading preview...

Model Overview

This model, miiikiik/Qwen3-8B-music-coa-sft, is an 8 billion parameter language model built upon the Qwen3 architecture. It has been specifically fine-tuned from a base model located at /mnt/tidal-alsh01/usr/zhongmeizhi/meiwangyi/models/Qwen3-8B-onto-fixed-traj.

Key Specialization

The primary differentiator for this model is its fine-tuning on the music_coa_sft dataset. This suggests a specialization in tasks or content related to music, potentially including music theory, composition, analysis, or other music-centric applications.

Training Details

The model was trained with the following key hyperparameters:

  • Learning Rate: 1e-05
  • Batch Size: 2 (train), 8 (eval) with 4 gradient accumulation steps, leading to a total effective batch size of 64.
  • Optimizer: ADAMW_TORCH
  • LR Scheduler: Cosine with 0.1 warmup steps
  • Epochs: 3.0

Intended Use

While specific intended uses and limitations require more information, its fine-tuning on a music-specific dataset indicates its suitability for applications within the music domain. Developers should consider its specialized training for tasks where music-related understanding or generation is crucial.