reasonwang/SkillGym-Qwen3.5-9B
SkillGym-Qwen3.5-9B is a 9 billion parameter language model developed by reasonwang, fine-tuned from Qwen3.5-9B with a 32768 token context length. It is specifically trained on SkillGym trajectories to enable agents to utilize skills, tools, and shell commands within a sandboxed environment. This model excels at agentic tasks requiring skill application, demonstrating improved performance on benchmarks like SkillGym, SkillEval, SkillsBench, and Skill-Use-Bench compared to its base model.
Loading preview...
SkillGym-Qwen3.5-9B Overview
SkillGym-Qwen3.5-9B is a 9 billion parameter model developed by reasonwang, specifically fine-tuned from Qwen3.5-9B to enhance agentic capabilities. This model is designed for agents to effectively use "skills"—defined as folders containing instructions, reference documents, and scripts—to solve tasks. It achieves this by applying relevant skills through tool calls or shell commands within a sandboxed workspace.
Key Capabilities
- Skill-Use Agent Policy: Optimized for tasks requiring an agent to consult and apply predefined skills.
- Tool and Shell Command Execution: Capable of executing tool calls and shell commands in a sandboxed environment.
- Enhanced Performance: Demonstrates significant improvements over its base model on agentic benchmarks, including SkillGym (59.5 vs 41.3), SkillEval (74.7 vs 65.7), SkillsBench (22.4 vs 14.8), and Skill-Use-Bench (49.6 vs 14.8).
- Reasoning Traces: Trained with reasoning traces to improve task-solving logic.
Good For
- Developing Skill-Based Agents: Ideal for creating agents that need to leverage specific, pre-defined skills and tools to complete complex tasks.
- Automated Task Execution: Suitable for scenarios where agents need to interact with environments via shell commands or tool APIs.
- Research in Agentic AI: Valuable for researchers exploring skill acquisition and application in large language models, particularly within the context of the SkillGym framework.
Limitations
- Not a General Chat Assistant: This model is specialized for agentic tasks and not tuned for general conversational use.
- Context Budget: Long task episodes may exhaust the 32768 token context length before completion.
- Repetitive Reasoning: Reasoning can occasionally become repetitive or hit per-turn token limits.