zhanxing/afm-crag-movie-sft-v2
The zhanxing/afm-crag-movie-sft-v2 model is an 8 billion parameter language model fine-tuned from Qwen/Qwen3-8B. It was specifically trained on the afm_movie_all_trajectories_v2_include_max_turns_train dataset, achieving a final validation loss of 0.0388. This model is optimized for tasks related to movie trajectories, leveraging its Qwen3-8B base for specialized performance in this domain.
Loading preview...
Overview
This model, afm-crag-movie-sft-v2, is an 8 billion parameter language model derived from the Qwen3-8B architecture. It has been fine-tuned on a specific dataset, afm_movie_all_trajectories_v2_include_max_turns_train, indicating a specialization in movie-related conversational or trajectory-based tasks. The training process involved 2 epochs with a learning rate of 1e-05 and a total batch size of 30, resulting in a low final validation loss of 0.0388.
Key Characteristics
- Base Model: Fine-tuned from Qwen/Qwen3-8B.
- Parameter Count: 8 billion parameters.
- Context Length: Supports a context length of 32768 tokens.
- Training Data: Specialized on the
afm_movie_all_trajectories_v2_include_max_turns_traindataset. - Performance: Achieved a validation loss of 0.0388, suggesting effective fine-tuning for its target domain.
Potential Use Cases
Given its specialized training, this model is likely suitable for applications requiring an understanding or generation of content related to movie trajectories or similar domain-specific interactions. Developers might consider it for:
- Movie recommendation systems.
- Dialogue generation for movie-related scenarios.
- Analysis of user interactions within movie databases or platforms.