Agent-Ark/Toucan-Qwen2.5-14B-Instruct-v0.1

TEXT GENERATIONConcurrent Unit Cost:1Model Size:14.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 2, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Agent-Ark/Toucan-Qwen2.5-14B-Instruct-v0.1 is a 14.8 billion parameter instruction-tuned causal language model, fine-tuned by Agent-Ark based on the Qwen2.5 architecture. It is specifically trained on the Toucan-1.5M dataset, the largest fully synthetic tool-agent dataset, to enhance tool-use capabilities in agentic LLMs. This model excels at complex multi-round, multi-turn, sequential, and parallel tool calls, outperforming larger models on tool-use benchmarks.

Loading preview...

Overview

Agent-Ark/Toucan-Qwen2.5-14B-Instruct-v0.1 is a 14.8 billion parameter model, fine-tuned from the Qwen2.5-14B-Instruct base model. Its primary distinction lies in its training on the Toucan-1.5M dataset, a comprehensive, fully synthetic tool-agent dataset comprising over 1.5 million trajectories. This dataset is derived from 495 real-world Model Context Protocols (MCPs) and involves over 2,000 tools, focusing on diverse and challenging tasks requiring multi-round, multi-turn, sequential, and parallel tool executions.

Key Capabilities

  • Advanced Tool Use: Specifically optimized for complex tool interaction, enabling LLMs to effectively utilize multiple tools in intricate scenarios.
  • Enhanced Agentic Performance: Models fine-tuned on Toucan-1.5M demonstrate superior performance on benchmarks like BFCL V3 and MCP-Universe, often surpassing larger, closed-source counterparts.
  • Diverse Task Handling: Trained on a rich variety of instances, including those focusing on irrelevance, diversification, and multi-turn interactions, ensuring robust performance across different tool-use patterns.

Good For

  • Developing agentic LLMs that require sophisticated tool-use capabilities.
  • Applications demanding multi-step reasoning and interaction with external tools or APIs.
  • Scenarios where smaller models need to achieve high performance in tool-augmented tasks, extending the Pareto frontier on relevant benchmarks.

For more technical details, refer to the Technical Report.