minsu0567/IAD-X1-SFT-si-answer-last
The minsu0567/IAD-X1-SFT-si-answer-last model is a 4.5 billion parameter language model fine-tuned from Qwen/Qwen3.5-4B. It was specifically trained on the PA_SFT_2_reordered_si_answer_last dataset, suggesting an optimization for specific question-answering or instruction-following tasks where the answer is expected at the end. This model is suitable for applications requiring a compact yet capable language model with a 32K context window, potentially excelling in tasks aligned with its specialized fine-tuning data.
Loading preview...
Model Overview
The minsu0567/IAD-X1-SFT-si-answer-last is a 4.5 billion parameter language model derived from the Qwen/Qwen3.5-4B base architecture. This model has undergone supervised fine-tuning (SFT) using the PA_SFT_2_reordered_si_answer_last dataset. The specific nature of this dataset, particularly the 'answer_last' indication, suggests that the model is optimized for scenarios where the desired output or answer is positioned at the end of a sequence or prompt.
Training Details
The fine-tuning process utilized the following key hyperparameters:
- Learning Rate: 1e-05
- Batch Size: 1 (train), 8 (eval)
- Gradient Accumulation Steps: 2, resulting in a total effective batch size of 2
- Optimizer: AdamW with specific betas and epsilon
- LR Scheduler: Cosine with 100 warmup steps
- Epochs: 1.0
This configuration indicates a focused, single-epoch fine-tuning approach designed to adapt the base Qwen3.5-4B model to the specific patterns present in the PA_SFT_2_reordered_si_answer_last dataset. The model retains the 32768 token context length of its base, making it suitable for tasks requiring processing of longer inputs while focusing on end-of-sequence responses.