KickItLikeShika/qwen-2.5-7b-taid-tooluse

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 19, 2026Architecture:Transformer Featherless Exclusive Cold

KickItLikeShika/qwen-2.5-7b-taid-tooluse is a 7.6 billion parameter language model based on the Qwen2.5 architecture, fine-tuned using Temporally Adaptive Interpolated Distillation (TAID). This model specializes in tool use, having been trained on a dedicated tool-use dataset. It achieves a 62.8% score on its evaluation set, making it suitable for applications requiring function calling and external tool interaction.

Loading preview...

Model Overview

KickItLikeShika/qwen-2.5-7b-taid-tooluse is a 7.6 billion parameter model built upon the Qwen2.5-7B-Instruct architecture. Its key differentiator is its training methodology: Temporally Adaptive Interpolated Distillation (TAID). This technique uses a larger Qwen2.5-32B-Instruct model as a teacher to distill knowledge into the smaller 7.6B parameter student model.

Key Capabilities

  • Tool Use Specialization: The model is specifically trained on a tool-use dataset, enhancing its ability to understand and execute function calls or interact with external tools.
  • Performance: Achieved a 62.8% score on its evaluation set, indicating proficiency in tool-use scenarios.
  • Efficient Knowledge Transfer: Leverages TAID for effective knowledge transfer from a larger, more capable teacher model, aiming for strong performance in a smaller footprint.

Use Cases

This model is particularly well-suited for applications requiring:

  • Function Calling: Integrating with APIs or external services where the model needs to generate structured calls.
  • Automated Workflows: Tasks that involve using tools to retrieve information, perform actions, or interact with systems.
  • Agentic AI Systems: Developing AI agents that can plan and execute multi-step tasks by utilizing various tools.