ipfipfipf/Qwen3.5-9B-sdpo-react-mathcodesearch-grpo-arm-e-step29
The ipfipfipf/Qwen3.5-9B-sdpo-react-mathcodesearch-grpo-arm-e-step29 is a 9 billion parameter language model based on Qwen/Qwen3.5-9B, specifically post-trained using GRPO and SDPO. It is optimized for multi-turn, native-tool-calling tasks involving math, code, and search, demonstrating strong performance on benchmarks like AIME, AMO-Bench, and OJBench. This model excels in complex reasoning and tool-use scenarios, making it suitable for applications requiring advanced problem-solving capabilities.
Loading preview...
Model Overview
This model, ipfipfipf/Qwen3.5-9B-sdpo-react-mathcodesearch-grpo-arm-e-step29, is a 9 billion parameter language model derived from Qwen/Qwen3.5-9B. It has undergone extensive RL post-training using a combination of GRPO and SDPO (self-skill objective) techniques. The training focused on a diverse mixture of math, code, and search tasks, specifically designed for multi-turn, native-tool-calling (ReAct-style) interactions.
Key Capabilities
- Enhanced Reasoning: Demonstrates strong performance on challenging benchmarks such as AIME 2024/2025 (up to 100% pass@8), AMO-Bench (28.0% pass@1), and OJBench (31.8% pass@1), indicating advanced problem-solving abilities.
- Native Tool Calling: Optimized for integrating external tools, with a significantly higher tool-use rate (84.3% on OJBench) compared to the base model, and supports XML-style
<function=...>calls. - Multi-turn Interactions: Trained on rollouts with up to 20 turns, enabling complex, iterative problem-solving.
- Language Model Only: While based on a vision-capable model, this checkpoint specifically focuses on the language model, with the vision tower and MTP heads removed during conversion to optimize for text-based tasks.
Good For
- Complex Math Problems: Excels in mathematical reasoning, as evidenced by high scores on AIME benchmarks.
- Code Generation and Debugging: Strong performance on coding benchmarks like OJBench suggests proficiency in code-related tasks.
- Search and Information Retrieval: Designed to effectively utilize search tools within multi-turn conversations.
- Applications Requiring Tool Use: Ideal for scenarios where the LLM needs to interact with external APIs or tools to complete tasks.
- Research in RLHF and Tool-Augmented LLMs: Represents a specific checkpoint (step 29) from a detailed training curve, offering insights into RL post-training methodologies.