code-critic-model/Qwen3-4B-Critic-SFT-DPO
Qwen3-4B-Critic-SFT-DPO is a 4 billion parameter critic model developed by code-critic-model, fine-tuned with Direct Preference Optimization (DPO) for steering large code agents. Building on Qwen3-4B-Critic-SFT, it provides structured critiques including error categories, evidence, and recovery actions. This model excels at improving the resolve rate of various coding agents on tasks like SWE-bench Verified by offering guidance rather than directly solving problems.
Loading preview...
Overview
This model, Qwen3-4B-Critic-SFT-DPO, is a 4 billion parameter critic model developed by code-critic-model. It is an enhanced version of Qwen3-4B-Critic-SFT, further trained using Direct Preference Optimization (DPO) on pairs of critiques. Its primary function is to act as a critic alongside a frozen coding agent, providing short, structured critiques at regular intervals. These critiques include detected error categories, supporting evidence, suggested recovery actions, task status, and overall guidance, effectively steering the agent's trajectory without directly writing code patches.
Key Capabilities
- Agent Steering: Provides actionable feedback to large code agents, improving their performance on complex coding tasks.
- DPO Enhancement: Utilizes DPO training on 1,409 preference pairs, where Claude Opus 4.6 judged critiques for correctness and clarity.
- Performance Improvement: Demonstrates improved resolve rates on the SWE-bench Verified dataset across various coding agents (e.g., Qwen3-32B, GPT-OSS-20B, GLM-4.7-Flash-30B-A3B) compared to its SFT-only predecessor.
- Structured Critiques: Generates critiques with specific error categories, evidence, and recovery actions.
Use Cases
This model is ideal for integrating into automated code development pipelines where a small, efficient critic can significantly enhance the performance of larger, more general-purpose coding agents. It is particularly suited for scenarios requiring iterative refinement and error correction in code generation and problem-solving.