Love2DL/qwne3_4B_shortSFT
Love2DL/qwne3_4B_shortSFT is a 4 billion parameter language model fine-tuned from Qwen/Qwen3-4B. This model was specifically trained on the sft_data dataset, indicating a focus on supervised fine-tuning tasks. With a context length of 32768 tokens, it is designed for applications requiring processing of longer sequences. Its primary strength lies in adapting the Qwen3-4B base model for specific instruction-following or task-oriented generation based on its training data.
Loading preview...
Model Overview
Love2DL/qwne3_4B_shortSFT is a 4 billion parameter language model derived from the Qwen/Qwen3-4B architecture. This model has undergone supervised fine-tuning (SFT) using the sft_data dataset, suggesting an optimization for specific instruction-following or task-oriented generation capabilities. It leverages a substantial context window of 32768 tokens, making it suitable for processing and generating longer text sequences.
Training Details
The model was trained with the following key hyperparameters:
- Learning Rate: 2e-05
- Batch Size: 2 (train), 8 (eval)
- Gradient Accumulation Steps: 12, leading to a total train batch size of 96
- Optimizer: AdamW with default betas and epsilon
- Scheduler: Cosine learning rate scheduler with a 0.1 warmup ratio
- Epochs: 3.0
This configuration indicates a focused fine-tuning approach to adapt the base Qwen3-4B model to the specific characteristics of the sft_data dataset.
Intended Use Cases
While specific intended uses are not detailed, models fine-tuned on SFT datasets are typically well-suited for:
- Instruction Following: Generating responses based on explicit instructions.
- Task-Specific Generation: Performing particular text generation tasks learned from the SFT data.
- Conversational AI: Engaging in dialogue where responses are guided by provided examples.
Users should evaluate its performance on their specific tasks, especially given the general nature of the sft_data dataset name.