suryadv/IncorrectTraceSFT-Qwen2.5-7B-AIME2025-incorrect
IncorrectTraceSFT-Qwen2.5-7B-AIME2025-incorrect is a 7.6 billion parameter Qwen2.5-7B model fine-tuned by suryadv. This model is specifically trained using 'incorrect traces' from the AIME2025 dataset, focusing on research into how such traces impact supervised fine-tuning. It is designed for research in mathematical reasoning and the study of SFT data selection strategies.
Loading preview...
Model Overview
The suryadv/IncorrectTraceSFT-Qwen2.5-7B-AIME2025-incorrect is a 7.6 billion parameter language model based on the Qwen2.5-7B architecture. It was developed by suryadv as part of research into the impact of incorrect traces in supervised fine-tuning (SFT), specifically for the AIME2025 dataset. This particular checkpoint, aime25_incorrect_3000, represents a condition from the "aime years" experiment.
Key Characteristics
- Base Model: Fine-tuned from
Qwen/Qwen2.5-7B. - Training Methodology: Utilizes SLIME with Megatron-LM, BF16, and AdamW, undergoing 188 optimizer updates.
- Data Focus: Trained using a dataset of "incorrect traces" related to mathematical reasoning problems from the AIME2025 competition.
- Context Length: Supports a context length of 32768 tokens.
Intended Use
This model is primarily released for research purposes in the following areas:
- Investigating mathematical reasoning capabilities of LLMs.
- Studying the effects and optimal strategies for using incorrect traces in supervised fine-tuning data selection.
It is important to note that while the training data may contain correct final answers, the intermediate reasoning steps were not necessarily verified. The model's performance should be evaluated using the provided evaluation scripts to understand its pass@1 and Monte Carlo standard error.