cuong1692001/Terminal-terminal_traj
The cuong1692001/Terminal-terminal_traj model is an 8 billion parameter language model fine-tuned from Qwen/Qwen3-8B. This model is specifically adapted for tasks related to the 'terminal_traj' dataset. It leverages the Qwen3 architecture, offering a substantial context length of 32768 tokens. Its primary differentiation lies in its specialized fine-tuning for specific trajectory-related applications.
Loading preview...
Model Overview
cuong1692001/Terminal-terminal_traj is an 8 billion parameter language model, fine-tuned from the robust Qwen/Qwen3-8B base model. This specialization targets tasks associated with the terminal_traj dataset, indicating an optimization for specific trajectory-related data processing or generation.
Key Characteristics
- Base Model: Qwen/Qwen3-8B
- Parameter Count: 8 billion parameters
- Context Length: Supports a substantial context window of 32768 tokens, inherited from its base architecture.
- Fine-tuning Focus: Specifically fine-tuned on the
terminal_trajdataset, suggesting enhanced performance for tasks within this domain.
Training Details
The model was trained with a learning rate of 1e-05, using an AdamW optimizer and a cosine learning rate scheduler over 2 epochs. The training utilized a distributed setup across 4 devices with a total batch size of 8.
Intended Use Cases
Given its fine-tuning on the terminal_traj dataset, this model is likely best suited for applications requiring specialized understanding or generation related to terminal trajectories. Users should consider its specific training data when evaluating its applicability for their particular use case.