Minbyul/AgentMercury-Qwen3.5-4B
Minbyul/AgentMercury-Qwen3.5-4B is a 4.5 billion parameter multimodal Qwen3.5-4B checkpoint, post-trained using agentic reinforcement learning on Model-Context-Protocol (MCP) tool-use environments. This model is specifically optimized for complex multi-turn agent tasks, excelling in tool correctness and environment state verification. It demonstrates significant improvements over its base model in agentic tool-use, competition math, and code benchmarks, making it suitable for applications requiring robust agentic capabilities and problem-solving.
Loading preview...
AgentMercury-Qwen3.5-4B Overview
AgentMercury-Qwen3.5-4B is a 4.5 billion parameter multimodal model based on the Qwen3.5-4B architecture, developed by Minbyul. Its core differentiator is its post-training with agentic reinforcement learning (RL) on Model-Context-Protocol (MCP) tool-use environments. This RL objective rewards the successful completion of real multi-turn agent tasks, focusing on correct tool calls and accurate final environment states, rather than just text generation.
Key Capabilities & Differentiators
- Agentic Reinforcement Learning: Trained using on-policy GRPO with 200-step MCP agentic RL, optimizing for practical agent task completion.
- Multimodal Base: Built upon Qwen3.5-4B, providing both text and vision capabilities.
- Clean-Minimum Checkpoint: This release represents the optimal training step where reward peaks without introducing degenerate generation or truncation, ensuring high-quality outputs.
- Broad Training Diversity: The training set includes approximately 2,300 agent environments, covering 63% of industries and 76% of tools from the source corpus.
Performance Highlights
AgentMercury-Qwen3.5-4B shows notable improvements over its base model in several key areas:
- Agentic / Tool-use: Significant gains on benchmarks like BFCL (+1.58) and τ³-bench (+0.041).
- Math & Reasoning: Improved performance in competition math benchmarks such as AIME 2026 (+0.094) and HMMT 2026-02 (+0.071).
- Code Generation: Enhanced capabilities in coding, with LiveCodeBench (v5+v6) showing a +0.069 improvement.
- Writing: Modest gains on WritingBench (+0.075).
When to Use This Model
This model is particularly well-suited for applications requiring:
- Robust Agentic Behavior: Ideal for tasks involving complex tool use, multi-step reasoning, and interaction with external environments.
- Problem Solving: Excels in mathematical and coding challenges where precise, verifiable solutions are needed.
- Multimodal Agent Applications: Leveraging its Qwen3.5-4B base, it can handle tasks that integrate both text and visual information within an agentic framework.