SeongryongJung/qwen3-8b-tooluse-grpo

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 2, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

SeongryongJung/qwen3-8b-tooluse-grpo is an 8 billion parameter language model fine-tuned from Qwen/Qwen3-8B. It utilizes the GRPO method specifically on a tool-use dataset, demonstrating a validation performance of 66.45% mean@16 on tool-use tasks. This model is optimized for applications requiring effective tool utilization and function calling capabilities.

Loading preview...

Model Overview

SeongryongJung/qwen3-8b-tooluse-grpo is an 8 billion parameter model derived from the Qwen3-8B architecture. It has been fine-tuned using the GRPO (Generalized Reinforcement Learning with Policy Optimization) method, specifically targeting tool-use capabilities.

Key Capabilities

  • Enhanced Tool Use: The model is specialized in understanding and executing tool-use instructions, as evidenced by its training on a dedicated tooluse dataset.
  • Performance Metrics: Achieved a peak validation performance of 66.45% on the val-aux/tooluse/reward/mean@16 metric during its 100-step fine-tuning process.
  • GRPO Fine-tuning: Leverages GRPO for improved performance in complex, multi-step reasoning and tool interaction scenarios.

Intended Use Cases

This model is particularly well-suited for applications that require:

  • Function Calling: Integrating with external APIs or functions based on natural language prompts.
  • Agentic Workflows: Developing AI agents that can interact with tools to accomplish tasks.
  • Complex Instruction Following: Handling instructions that necessitate the use of specific tools or external knowledge sources.