shaffhausenquant/OpenThinkerAgent-32B-SFT-3.16K

TEXT GENERATIONPricing:Input $0.408 / Cached $0.0816 / Output $1.972Concurrent Unit Cost:2Model Size:32BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 11, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

OpenThinkerAgent-32B-SFT-3.16K by shaffhausenquant is a 32 billion parameter language model post-trained from Qwen3-32B. It is fine-tuned using the OpenThoughts-Agent-SFT-3.16K dataset, specifically optimized for agentic tasks and complex problem-solving. This model excels in benchmarks like SWE-Bench-Verified-100 and Terminal-Bench 2.0, making it suitable for applications requiring robust agentic capabilities.

Loading preview...

OpenThinkerAgent-32B-SFT-3.16K: An Agentic Language Model

OpenThinkerAgent-32B-SFT-3.16K is a 32 billion parameter model developed by shaffhausenquant, fine-tuned from the Qwen3-32B base model. This model is part of the OpenThoughts-Agent initiative, which focuses on curating high-quality datasets for training agentic models.

Key Capabilities & Training

This model is specifically designed for agentic tasks, leveraging a full-parameter Supervised Fine-Tuning (SFT) approach on the 3,160-example OpenThoughts-Agent-SFT-3.16K dataset. The dataset comprises (task, agent-trajectory) pairs from diverse sources like SWE-Smith and StackExchange, with trajectories generated by GLM-4.7-AWQ and filtered for at least 5 model turns. The training utilized a cutoff length of 32768 tokens, bf16 precision, and DeepSpeed ZeRO-3.

Performance Highlights

Evaluated in the terminus-2 harness, OpenThinkerAgent-32B-SFT-3.16K demonstrates improved performance over its base model, Qwen3-32B, across key agentic benchmarks:

  • SWE-Bench-Verified-100: Achieved 30.3 (vs. 26.7 for Qwen3-32B)
  • OpenThoughts-TBLite: Achieved 24.7 (vs. 13.7 for Qwen3-32B)
  • Terminal-Bench 2.0: Achieved 16.9 (vs. 7.5 for Qwen3-32B)

Ideal Use Cases

This model is particularly well-suited for applications requiring advanced agentic reasoning and problem-solving, such as:

  • Automated code generation and debugging (e.g., SWE-Bench tasks)
  • Complex task execution in terminal environments
  • Developing intelligent agents capable of multi-turn interactions and planning