code-critic-model/Qwen3-4B-Critic-SFT
The code-critic-model/Qwen3-4B-Critic-SFT is a 4 billion parameter Qwen3-based critic model, fine-tuned by Shubham Gandhi, Yiqing Xie, Atharva Naik, Ruichen Zhu, and Carolyn Rose, designed to steer large code agents. With a 32,768 token context length, it provides structured critiques on agent trajectories, identifying error categories, evidence, and recovery actions. This model excels at improving the resolve rate of various coding agents on tasks like SWE-bench Verified by offering guidance rather than directly solving problems.
Loading preview...
Overview
code-critic-model/Qwen3-4B-Critic-SFT is a 4 billion parameter critic model based on Qwen3-4B-Instruct-2507, developed by Shubham Gandhi et al. as part of the "Steer, Don't Solve: Training Small Critic Models for Large Code Agents" research. Unlike traditional code generation models, this model acts as a critic, providing structured feedback to a separate, frozen coding agent. It analyzes the agent's trajectory and outputs a critique including error categories, evidence, recovery actions, and overall guidance, aiming to steer the agent towards a solution rather than generating code itself.
Key Capabilities
- Trajectory Analysis: Reads agent trajectories and identifies issues.
- Structured Critiques: Provides feedback in a structured format, detailing error categories, evidence, and recovery actions.
- Agent Steering: Improves the performance of various coding agents by guiding them through complex tasks.
- High Context Length: Supports a sequence length of 32,768 tokens, allowing for analysis of extensive agent interactions.
Performance
Evaluated on SWE-bench Verified, the model significantly boosts the resolve rate of several coding agents. For instance, it improved Qwen3-32B's resolve rate from 10.2% to 11.4% and GPT-OSS-120B's from 16.8% to 31.2% when used as a critic. This SFT version serves as the starting point for the DPO critic Qwen3-4B-Critic-SFT-DPO.
Training Details
The model was fine-tuned using LLaMA-Factory on a dataset of 6,447 examples from R2E-Gym and SWE-bench Verified tasks. The training data consisted of trajectories generated by CWM-32B and Qwen3-Next-80B-A3B-Instruct agents, with critiques provided by Claude Opus 4.6.