Minbyul/AgentMercury-Qwen3.5-4B

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 14, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Minbyul/AgentMercury-Qwen3.5-4B is a 4.5 billion parameter multimodal Qwen3.5-4B checkpoint, post-trained using agentic reinforcement learning on Model-Context-Protocol (MCP) tool-use environments. This model is specifically optimized for complex multi-turn agent tasks, excelling in tool correctness and environment state verification. It demonstrates significant improvements over its base model in agentic tool-use, competition math, and code benchmarks, making it suitable for applications requiring robust agentic capabilities and problem-solving.

Loading preview...

AgentMercury-Qwen3.5-4B Overview

AgentMercury-Qwen3.5-4B is a 4.5 billion parameter multimodal model based on the Qwen3.5-4B architecture, developed by Minbyul. Its core differentiator is its post-training with agentic reinforcement learning (RL) on Model-Context-Protocol (MCP) tool-use environments. This RL objective rewards the successful completion of real multi-turn agent tasks, focusing on correct tool calls and accurate final environment states, rather than just text generation.

Key Capabilities & Differentiators

  • Agentic Reinforcement Learning: Trained using on-policy GRPO with 200-step MCP agentic RL, optimizing for practical agent task completion.
  • Multimodal Base: Built upon Qwen3.5-4B, providing both text and vision capabilities.
  • Clean-Minimum Checkpoint: This release represents the optimal training step where reward peaks without introducing degenerate generation or truncation, ensuring high-quality outputs.
  • Broad Training Diversity: The training set includes approximately 2,300 agent environments, covering 63% of industries and 76% of tools from the source corpus.

Performance Highlights

AgentMercury-Qwen3.5-4B shows notable improvements over its base model in several key areas:

  • Agentic / Tool-use: Significant gains on benchmarks like BFCL (+1.58) and τ³-bench (+0.041).
  • Math & Reasoning: Improved performance in competition math benchmarks such as AIME 2026 (+0.094) and HMMT 2026-02 (+0.071).
  • Code Generation: Enhanced capabilities in coding, with LiveCodeBench (v5+v6) showing a +0.069 improvement.
  • Writing: Modest gains on WritingBench (+0.075).

When to Use This Model

This model is particularly well-suited for applications requiring:

  • Robust Agentic Behavior: Ideal for tasks involving complex tool use, multi-step reasoning, and interaction with external environments.
  • Problem Solving: Excels in mathematical and coding challenges where precise, verifiable solutions are needed.
  • Multimodal Agent Applications: Leveraging its Qwen3.5-4B base, it can handle tasks that integrate both text and visual information within an agentic framework.