luganoquant/OpenThinkerAgent-32B
OpenThinkerAgent-32B is a 32 billion parameter language model developed by OpenThoughts-Agent, post-trained from Qwen3-32B. It is specifically fine-tuned on the 100,000-example OpenThoughts-Agent-SFT-100K dataset for agentic tasks. This model excels in complex agentic benchmarks, demonstrating strong performance in areas like code generation, terminal interaction, and problem-solving across various domains. It is designed for applications requiring advanced autonomous agent capabilities.
Loading preview...
OpenThinkerAgent-32B: An Agentic Language Model
OpenThinkerAgent-32B is a 32 billion parameter model developed by OpenThoughts-Agent, built upon the Qwen3-32B architecture. This model is specifically designed for agentic tasks, having undergone full-parameter supervised fine-tuning (SFT) on the extensive OpenThoughts-Agent-SFT-100K dataset. This dataset comprises 100,000 examples of task-agent trajectory pairs, sourced from diverse origins like SWE-Smith, StackExchange, and IssueTasks, with trajectories generated by GLM-4.7-AWQ and filtered for traces with at least five model turns.
Key Capabilities & Performance
OpenThinkerAgent-32B demonstrates significant performance improvements over its base model, Qwen3-32B, particularly in agentic benchmarks. Evaluated in the terminus-2 harness, it achieves:
- SWE-Bench-Verified-100: 55.7% (compared to 26.7% for Qwen3-32B)
- OpenThoughts-TBLite: 41.3% (compared to 13.7% for Qwen3-32B)
- Terminal-Bench 2.0: 26.2% (compared to 7.5% for Qwen3-32B)
Across a suite of seven agentic benchmarks, OpenThinkerAgent-32B achieves an average accuracy of 44.8%, positioning it as the strongest open-data 32B model for agentic tasks. Its training involved a learning rate of 4e-05, cosine scheduler, and a global batch size of 96 over 5 epochs, with a cutoff length of 32768 tokens.
Ideal Use Cases
- Autonomous Agents: Developing AI agents capable of complex problem-solving and interaction.
- Code Generation & Debugging: Tasks requiring high proficiency in understanding and generating code, as evidenced by its SWE-Bench performance.
- Terminal Interaction: Applications that involve navigating and executing commands within terminal environments.
- Multi-turn Reasoning: Scenarios demanding sustained, multi-step reasoning and decision-making.