suryadv/IncorrectTraceSFT-Qwen2.5-7B-AIME-SplitB-incorrect

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 15, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The suryadv/IncorrectTraceSFT-Qwen2.5-7B-AIME-SplitB-incorrect is a 7.6 billion parameter Qwen2.5-7B model, fine-tuned by suryadv, specifically from the "held_out_incorrect_6000" condition of the AIME splits experiment. This model is designed for research into how incorrect traces influence supervised fine-tuning, particularly in mathematical reasoning tasks. It focuses on understanding the impact of unverified intermediate reasoning steps in training data on model performance.

Loading preview...

Overview

This model, IncorrectTraceSFT-Qwen2.5-7B-AIME-SplitB-incorrect, is a 7.6 billion parameter Qwen2.5-7B variant developed by suryadv. It is a full fine-tune derived from research into "How Should Incorrect Traces Be Used in Supervised Fine-Tuning?" Specifically, it represents the held_out_incorrect_6000 condition within the AIME splits experiment.

Key Characteristics

  • Base Model: Qwen2.5-7B, revision d149729398750b98c0af14eb82c78cfe92750796.
  • Training Method: Utilizes SLIME with Megatron-LM, BF16, a global batch size of 64, and AdamW with a peak learning rate of 5e-6.
  • Focus: Investigates the role of incorrect traces in supervised fine-tuning, particularly in the context of mathematical reasoning.
  • Data: Trained using the suryadv/qwen3-8b-aime-2009-2024-16x dataset.

Intended Use

This model is primarily released for research purposes concerning mathematical reasoning and the selection of data for Supervised Fine-Tuning (SFT). It's important to note that while the training data may contain correct final answers, the intermediate reasoning steps were not necessarily verified. The study's ID/OOD (in-distribution/out-of-distribution) refers to problem identity relative to SFT, not pretraining exposure. Evaluation can be performed using the provided incorrect_trace_sft.evaluate script to report pass@1 and Monte Carlo standard error.