zhanxing/ours-crag-movie-sft-v2
The zhanxing/ours-crag-movie-sft-v2 is an 8 billion parameter language model fine-tuned from Qwen/Qwen3-8B. This model has been specialized using the ours_crag_movie_sft_v2_train dataset, achieving a final validation loss of 0.0837. It is designed for tasks related to its specific fine-tuning domain, likely movie-related content generation or analysis, leveraging a 32768 token context length.
Loading preview...
Model Overview
The zhanxing/ours-crag-movie-sft-v2 is an 8 billion parameter language model built upon the Qwen3-8B architecture. It has been specifically fine-tuned on the ours_crag_movie_sft_v2_train dataset, indicating a specialization in a particular domain, likely related to movie content.
Training Details
The model underwent training for 2 epochs with a learning rate of 1e-05 and a total batch size of 30, utilizing a multi-GPU setup. The training process used an AdamW optimizer with a cosine learning rate scheduler and a warmup ratio of 0.1. Over the course of training, the validation loss steadily decreased, reaching a final reported loss of 0.0837.
Key Characteristics
- Base Model: Qwen/Qwen3-8B
- Parameter Count: 8 Billion
- Context Length: 32768 tokens
- Fine-tuning Dataset:
ours_crag_movie_sft_v2_train - Achieved Validation Loss: 0.0837
Potential Use Cases
Given its fine-tuning on a specific dataset, this model is likely best suited for applications that align with the nature of the ours_crag_movie_sft_v2_train data. This could include tasks such as movie plot generation, script analysis, character dialogue creation, or other movie-centric natural language processing tasks.