shaffhausenquant/OpenThinkerAgent-32B-SFT-316

TEXT GENERATIONPricing:Input $0.408 / Cached $0.0816 / Output $1.972Concurrent Unit Cost:2Model Size:32BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 11, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

OpenThinkerAgent-32B-SFT-316 by shaffhausenquant is a 32 billion parameter language model, post-trained from Qwen3-32B. It is specifically fine-tuned for agentic tasks using the OpenThoughts-Agent-SFT-316 dataset, which comprises agent-trajectory pairs from various task sources. This model excels in agentic benchmarks like OpenThoughts-TBLite and Terminal-Bench 2.0, making it suitable for developing autonomous agents.

Loading preview...

OpenThinkerAgent-32B-SFT-316 Overview

OpenThinkerAgent-32B-SFT-316 is a 32 billion parameter model developed by shaffhausenquant, derived from the Qwen3-32B architecture. It is the result of a full-parameter Supervised Fine-Tuning (SFT) process using the specialized OpenThoughts-Agent-SFT-316 dataset. This dataset consists of 316 high-quality agent-trajectory examples, sourced from tasks like SWE-Smith, StackExchange-SuperUser, StackExchange-Tezos, and IssueTasks, with trajectories generated by GLM-4.7-AWQ and filtered for traces with at least 5 model turns.

Key Capabilities & Performance

This model is specifically optimized for agentic workloads, demonstrating strong performance in relevant benchmarks:

  • OpenThoughts-TBLite: Achieves 24.2 pass@1, significantly outperforming its base model, Qwen3-32B (13.7).
  • Terminal-Bench 2.0: Scores 13.1 pass@1, also surpassing Qwen3-32B (7.5).

Training involved a learning rate of 4e-05, cosine scheduler with 0.1 warmup ratio, global batch size of 96, and 7 epochs, utilizing bf16 precision and DeepSpeed ZeRO-3.

Ideal Use Cases

  • Developing autonomous agents: Its specialized training makes it highly suitable for tasks requiring agentic reasoning and execution.
  • Agentic research and development: Provides a strong foundation for exploring and building advanced agentic systems.
  • Benchmarking agent performance: Can serve as a robust baseline or target for evaluating new agentic approaches.