bielquants/OpenThinkerAgent-32B
OpenThinkerAgent-32B is a 32 billion parameter language model developed by OpenThoughts-Agent, post-trained from Qwen3-32B. It is fine-tuned with full-parameter SFT on the 100,000-example OpenThoughts-Agent-SFT-100K dataset, specifically designed for agentic tasks. This model is optimized for complex problem-solving and tool-use scenarios, achieving strong performance across various agentic benchmarks.
Loading preview...
OpenThinkerAgent-32B: An Agentic Language Model
OpenThinkerAgent-32B is a 32 billion parameter model developed by the OpenThoughts-Agent team, built upon the Qwen3-32B architecture. This model is specifically designed and optimized for agentic capabilities, focusing on complex task execution and problem-solving.
Key Capabilities and Training
The model's agentic prowess stems from its post-training process, which involves full-parameter Supervised Fine-Tuning (SFT) on the extensive OpenThoughts-Agent-SFT-100K dataset. This dataset comprises 100,000 (task, agent-trajectory) pairs sourced from diverse domains like SWE-Smith, StackExchange, and IssueTasks. The trajectories were generated by GLM-4.7-AWQ and filtered for traces with at least 5 model turns, ensuring high-quality, multi-step reasoning examples.
Performance Highlights
OpenThinkerAgent-32B demonstrates significant improvements over its base model, Qwen3-32B, across several agentic benchmarks. Evaluated in the terminus-2 harness, it shows substantial gains:
- SWE-Bench-Verified-100: 55.7% (vs. 26.7% for Qwen3-32B)
- OpenThoughts-TBLite: 41.3% (vs. 13.7% for Qwen3-32B)
- Terminal-Bench 2.0: 26.2% (vs. 7.5% for Qwen3-32B)
Across a suite of seven agentic benchmarks, OpenThinkerAgent-32B achieves an average accuracy of 44.8%, positioning it as a leading open-data model in the 32B parameter class for agentic tasks.
Use Cases
This model is particularly well-suited for applications requiring:
- Automated problem-solving: Tackling complex tasks that involve multiple steps and tool interactions.
- Code generation and debugging: Demonstrated by its strong performance on SWE-Bench.
- Terminal-based operations: Excelling in environments requiring command-line interaction.
- Agentic workflows: Developing AI agents capable of planning, executing, and refining actions.