KickItLikeShika/qwen2.5-7b-infosft-tooluse

TEXT GENERATIONConcurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 23, 2026Architecture:Transformer Featherless Exclusive Cold

KickItLikeShika/qwen2.5-7b-infosft-tooluse is a 7.6 billion parameter language model based on the Qwen2.5 architecture, fine-tuned for tool use capabilities. It leverages the InfoSFT (Information-Aware Token Weighting) method to enhance learning efficiency and was trained on a specialized tool-use dataset. This model is designed to excel in tasks requiring external tool integration and scored 67% on its evaluation set.

Loading preview...

Model Overview

KickItLikeShika/qwen2.5-7b-infosft-tooluse is a 7.6 billion parameter model built upon the Qwen2.5 architecture, specifically fine-tuned for tool-use applications. This model incorporates the InfoSFT (Information-Aware Token Weighting) method, which aims to improve learning by focusing on more informative tokens and reducing the impact of less relevant ones. The training utilized a dedicated tool-use dataset, originally released with the Self-Distillation Fine-tuning (SDFT) framework.

Key Capabilities

  • Enhanced Tool Use: Specifically trained on a tool-use dataset, making it proficient in scenarios requiring interaction with external functions or APIs.
  • InfoSFT Integration: Benefits from Information-Aware Token Weighting, a technique designed to optimize the learning process by emphasizing critical information.
  • Performance: Achieved a 67% score on its evaluation set, indicating its effectiveness in tool-use tasks.

Good For

  • Function Calling: Ideal for applications where the model needs to identify and call external tools or APIs based on user prompts.
  • Agentic Workflows: Suitable for building AI agents that can leverage various tools to accomplish complex tasks.
  • Research in Tool Learning: Provides a base for further experimentation and development in the field of language models interacting with tools.