miiikiik/Qwen3-8B-music-movie-coa-sft

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 15, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The miiikiik/Qwen3-8B-music-movie-coa-sft model is an 8 billion parameter language model, fine-tuned from a Qwen3-8B base model. It specializes in tasks related to music and movies, having undergone further supervised fine-tuning on the movie_coa_sft dataset. This model is designed for applications requiring nuanced understanding and generation within these specific entertainment domains, leveraging its 32768 token context length for comprehensive processing.

Loading preview...

Model Overview

The miiikiik/Qwen3-8B-music-movie-coa-sft is an 8 billion parameter language model, fine-tuned from an existing Qwen3-8B variant that was previously adapted for music-related tasks. This iteration has undergone additional supervised fine-tuning specifically on the movie_coa_sft dataset, indicating a specialization in movie-related content and potentially a combination of music and movie contexts.

Key Characteristics

  • Base Model: Qwen3-8B architecture.
  • Parameter Count: 8 billion parameters.
  • Context Length: Supports a substantial context window of 32768 tokens.
  • Specialization: Fine-tuned on the movie_coa_sft dataset, building upon prior music-related fine-tuning.

Training Details

The model was trained using the following hyperparameters:

  • Learning Rate: 1e-05
  • Batch Size: train_batch_size of 2, eval_batch_size of 8, with a gradient_accumulation_steps of 4, resulting in a total_train_batch_size of 64.
  • Optimizer: ADAMW_TORCH with default betas and epsilon.
  • Scheduler: Cosine learning rate scheduler with a 0.1 warmup ratio.
  • Epochs: 3.0 epochs.

Intended Use Cases

This model is particularly suited for applications requiring detailed understanding, generation, or analysis within the music and movie domains, especially those benefiting from its specific fine-tuning on movie-related conversational or analytical data. Its large context window allows for processing longer inputs and generating more coherent and contextually relevant outputs in these areas.