schwyzquants/OpenThinkerAgent-32B
OpenThinkerAgent-32B is a 32 billion parameter language model developed by OpenThoughts-Agent, post-trained from Qwen3-32B with a 32768 token context length. It is specifically fine-tuned for agentic tasks using the 100,000-example OpenThoughts-Agent-SFT-100K dataset. This model excels in agentic benchmarks, demonstrating strong performance across various problem-solving and interactive environments.
Loading preview...
OpenThinkerAgent-32B: An Agentic Language Model
OpenThinkerAgent-32B is a 32 billion parameter model developed by OpenThoughts-Agent, specifically engineered for agentic capabilities. It is post-trained from Qwen3-32B using a comprehensive 100,000-example supervised fine-tuning (SFT) dataset, OpenThoughts-Agent-SFT-100K, which consists of task and agent-trajectory pairs. The training data was generated by GLM-4.7-AWQ and filtered for traces with at least five model turns, ensuring high-quality agentic interactions.
Key Capabilities & Performance
- Agentic Task Proficiency: Demonstrates significant improvements over its base model, Qwen3-32B, in agentic benchmarks.
- Benchmark Leader: Achieves an average score of 44.8 across a seven-benchmark suite, making it the strongest open-data 32B model for agentic tasks.
- Specific Benchmark Gains: Shows substantial performance increases on benchmarks like SWE-Bench-Verified-100 (from 26.7 to 55.7), OpenThoughts-TBLite (from 13.7 to 41.3), and Terminal-Bench 2.0 (from 7.5 to 26.2).
- Extensive Context: Supports a context length of 32768 tokens, suitable for complex, multi-turn agentic interactions.
Training Details
The model was trained with a learning rate of 4e-05, a cosine scheduler with 0.1 warmup ratio, a global batch size of 96, and 5 epochs, utilizing bf16 precision and DeepSpeed ZeRO-3. The training dataset, OpenThoughts-Agent-SFT-100K, is curated from diverse sources including SWE-Smith and StackExchange.
Ideal Use Cases
- Automated Problem Solving: Excellent for tasks requiring sequential decision-making and interaction with environments.
- Code Generation & Debugging: Strong performance on benchmarks like SWE-Bench-Verified suggests utility in software engineering tasks.
- Interactive Agents: Suitable for developing agents that can engage in multi-turn dialogues and execute complex plans.