SeongryongJung/Qwen3-4B-Tooluse-RLSD-TR
The SeongryongJung/Qwen3-4B-Tooluse-RLSD-TR is a 4 billion parameter Qwen3-based language model specifically fine-tuned for tool-use capabilities. Utilizing the RLSD_TR method, this model demonstrates a best validation mean@16 score of 61.03% on tool-use tasks. It is optimized for applications requiring models to effectively interact with and utilize external tools, offering a compact yet capable solution for tool-augmented language generation.
Loading preview...
Overview
SeongryongJung/Qwen3-4B-Tooluse-RLSD-TR is a 4 billion parameter model built upon the Qwen3 architecture, specifically fine-tuned for enhanced tool-use capabilities. This model leverages the RLSD_TR (Reinforcement Learning with Self-Distillation and Trust-Region) method during its training, focusing on improving its ability to interact with and utilize external tools effectively. It was trained on a dedicated "Tool-use / tooluse" dataset with a batch size of 32 over 100 steps.
Key Capabilities
- Tool-Use Optimization: Achieved a best validation
mean@16score of 61.03% on tool-use specific tasks, indicating proficiency in tool interaction. - Qwen3 Base: Benefits from the foundational strengths of the Qwen3-4B model.
- RLSD_TR Fine-tuning: Utilizes an advanced reinforcement learning method for specialized performance in tool-use scenarios.
- Compact Size: At 4 billion parameters, it offers a relatively efficient solution for tool-augmented applications.
Good For
- Developing applications that require language models to integrate with external APIs or tools.
- Use cases demanding tool-augmented reasoning and action generation.
- Scenarios where a smaller, specialized model for tool interaction is preferred over larger, general-purpose models.