muradil211/ToolWeave_stage3

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 19, 2026Architecture:Transformer0.0K Featherless Exclusive Cold

muradil211/ToolWeave_stage3 is a 4 billion parameter language model based on the Qwen3-4B-Instruct family, specifically designed for advanced multi-turn tool-calling capabilities. This model utilizes Boundary-Guided Online Reinforcement Learning, incorporating verified online data synthesis and multi-turn progress rewards. It excels in complex tool-use scenarios by detecting capability boundaries and performing strict execution and semantic validation, making it suitable for developing robust AI agents that interact with external tools.

Loading preview...

ToolWeave Stage 3: Advanced Multi-Turn Tool-Calling

muradil211/ToolWeave_stage3 is the final release in the ToolWeave project's Stage 3, built upon the Qwen3-4B-Instruct base model. This 4 billion parameter model is specifically engineered for multi-turn tool-use learning through a novel approach called Boundary-Guided Online Reinforcement Learning.

Key Capabilities & Features

  • Boundary-Guided Learning: Enhances tool-use by detecting the operational boundaries of tools.
  • Verified Online Data Synthesis: Generates and validates training data in real-time to improve agent performance.
  • Strict Execution & Semantic Validation: Ensures tool calls are both syntactically correct and semantically appropriate for the given context.
  • Dynamic Replay: Utilizes past interactions to refine future tool-calling strategies.
  • Combined Global/Local Tool-Call Credit: Optimizes reward signals for more effective learning in complex sequences.
  • Multi-Turn Progress Reward: A specialized training signal designed to improve performance over extended tool-use dialogues.

Performance

Evaluated on a balanced 400-row held-in set, ToolWeave Stage 3 achieved an overall accuracy of 48.50% in complete-entry BFCL Multi-Turn tasks. This includes specific scores across different categories:

  • Base: 56.00%
  • Missing Function: 50.00%
  • Missing Parameter: 42.00%
  • Long Context: 46.00%

When to Use This Model

This model is ideal for developers building AI agents that require sophisticated, reliable, and multi-step interactions with external tools. Its specialized training in boundary detection and verified data synthesis makes it particularly strong for applications where tool-calling accuracy and robustness in complex, multi-turn scenarios are critical.