stgallenquants/OpenThinker-32B

TEXT GENERATIONPricing:Input $2.72 / Output $4.8Concurrent Unit Cost:2Model Size:32.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 28, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

OpenThinker-32B by stgallenquants is a 32.8 billion parameter instruction-tuned causal language model, fine-tuned from Qwen2.5-32B-Instruct. It is specifically optimized for advanced reasoning tasks, demonstrating strong performance on benchmarks like MATH500 and GPQA Diamond. This model excels in complex problem-solving and knowledge-intensive applications, leveraging its training on the OpenThoughts-114k dataset.

Loading preview...

Overview

OpenThinker-32B is a 32.8 billion parameter language model developed by stgallenquants, fine-tuned from the Qwen2.5-32B-Instruct base model. Its primary distinction comes from its training on the OpenThoughts-114k dataset, which is derived by distilling DeepSeek-R1. This process aims to enhance the model's reasoning capabilities.

Key Capabilities & Performance

OpenThinker-32B demonstrates strong performance in reasoning and knowledge-based tasks, as evaluated using the open-source tool Evalchemy. Key benchmark results include:

  • MATH500: 90.6%
  • GPQA Diamond: 61.6%
  • LCBv2: 68.9%

These scores indicate its proficiency in mathematical problem-solving and complex question answering, often outperforming other 32B models like LIMO-32B and s1.1-32B in specific reasoning benchmarks.

Training Details

The model was fine-tuned for 3 epochs with a 16k context length using LlamaFactory. The training utilized 8xH100 P5 nodes on AWS SageMaker, taking approximately 90 hours. The project emphasizes full transparency, providing access to its model weights, datasets, data generation code, and evaluation code.

Intended Uses

This model is well-suited for applications requiring advanced reasoning, mathematical understanding, and accurate knowledge retrieval. Its open-source nature and detailed training methodology make it a valuable resource for researchers and developers focusing on improving LLM reasoning.