ipfipfipf/Qwen3.5-9B-sdpo-react-mathcodesearch-grpo-arm-e-step29

VISIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 1, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The ipfipfipf/Qwen3.5-9B-sdpo-react-mathcodesearch-grpo-arm-e-step29 is a 9 billion parameter language model based on Qwen/Qwen3.5-9B, specifically post-trained using GRPO and SDPO. It is optimized for multi-turn, native-tool-calling tasks involving math, code, and search, demonstrating strong performance on benchmarks like AIME, AMO-Bench, and OJBench. This model excels in complex reasoning and tool-use scenarios, making it suitable for applications requiring advanced problem-solving capabilities.

Loading preview...

Model Overview

This model, ipfipfipf/Qwen3.5-9B-sdpo-react-mathcodesearch-grpo-arm-e-step29, is a 9 billion parameter language model derived from Qwen/Qwen3.5-9B. It has undergone extensive RL post-training using a combination of GRPO and SDPO (self-skill objective) techniques. The training focused on a diverse mixture of math, code, and search tasks, specifically designed for multi-turn, native-tool-calling (ReAct-style) interactions.

Key Capabilities

  • Enhanced Reasoning: Demonstrates strong performance on challenging benchmarks such as AIME 2024/2025 (up to 100% pass@8), AMO-Bench (28.0% pass@1), and OJBench (31.8% pass@1), indicating advanced problem-solving abilities.
  • Native Tool Calling: Optimized for integrating external tools, with a significantly higher tool-use rate (84.3% on OJBench) compared to the base model, and supports XML-style <function=...> calls.
  • Multi-turn Interactions: Trained on rollouts with up to 20 turns, enabling complex, iterative problem-solving.
  • Language Model Only: While based on a vision-capable model, this checkpoint specifically focuses on the language model, with the vision tower and MTP heads removed during conversion to optimize for text-based tasks.

Good For

  • Complex Math Problems: Excels in mathematical reasoning, as evidenced by high scores on AIME benchmarks.
  • Code Generation and Debugging: Strong performance on coding benchmarks like OJBench suggests proficiency in code-related tasks.
  • Search and Information Retrieval: Designed to effectively utilize search tools within multi-turn conversations.
  • Applications Requiring Tool Use: Ideal for scenarios where the LLM needs to interact with external APIs or tools to complete tasks.
  • Research in RLHF and Tool-Augmented LLMs: Represents a specific checkpoint (step 29) from a detailed training curve, offering insights into RL post-training methodologies.