zhanxing/ours-crag-movie-sft-v2

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 5, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The zhanxing/ours-crag-movie-sft-v2 is an 8 billion parameter language model fine-tuned from Qwen/Qwen3-8B. This model has been specialized using the ours_crag_movie_sft_v2_train dataset, achieving a final validation loss of 0.0837. It is designed for tasks related to its specific fine-tuning domain, likely movie-related content generation or analysis, leveraging a 32768 token context length.

Loading preview...

Model Overview

The zhanxing/ours-crag-movie-sft-v2 is an 8 billion parameter language model built upon the Qwen3-8B architecture. It has been specifically fine-tuned on the ours_crag_movie_sft_v2_train dataset, indicating a specialization in a particular domain, likely related to movie content.

Training Details

The model underwent training for 2 epochs with a learning rate of 1e-05 and a total batch size of 30, utilizing a multi-GPU setup. The training process used an AdamW optimizer with a cosine learning rate scheduler and a warmup ratio of 0.1. Over the course of training, the validation loss steadily decreased, reaching a final reported loss of 0.0837.

Key Characteristics

  • Base Model: Qwen/Qwen3-8B
  • Parameter Count: 8 Billion
  • Context Length: 32768 tokens
  • Fine-tuning Dataset: ours_crag_movie_sft_v2_train
  • Achieved Validation Loss: 0.0837

Potential Use Cases

Given its fine-tuning on a specific dataset, this model is likely best suited for applications that align with the nature of the ours_crag_movie_sft_v2_train data. This could include tasks such as movie plot generation, script analysis, character dialogue creation, or other movie-centric natural language processing tasks.