trinityomni/OpenThinker-32B
TEXT GENERATIONPricing:Input $2.72 / Output $4.8Concurrent Unit Cost:2Model Size:32.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 7, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold
OpenThinker-32B by trinityomni is a 32.8 billion parameter instruction-tuned language model, fine-tuned from Qwen2.5-32B-Instruct. It is specifically optimized for advanced reasoning and mathematical tasks, demonstrating strong performance on benchmarks like AIME, MATH, and GPQA Diamond. This model is designed for applications requiring high-level cognitive abilities and complex problem-solving.
Loading preview...
OpenThinker-32B: A Reasoning-Optimized LLM
OpenThinker-32B is a 32.8 billion parameter language model developed by trinityomni, fine-tuned from Qwen/Qwen2.5-32B-Instruct. Its core differentiation lies in its optimization for complex reasoning and mathematical problem-solving, achieved through fine-tuning on the specialized OpenThoughts-114k dataset.
Key Capabilities & Performance
- Enhanced Reasoning: Demonstrates strong performance across various reasoning benchmarks, including AIME24 I/II (66.0), AIME25 I (53.3), and GPQA Diamond (61.6).
- Mathematical Proficiency: Achieves a score of 90.6 on the MATH500 benchmark, indicating robust mathematical problem-solving abilities.
- Open-Source Ecosystem: The model's weights, datasets, data generation code, evaluation code (Evalchemy), and training code are all publicly available, fostering transparency and community contributions.
- Training Details: Fine-tuned for 3 epochs with a 16k context length using LlamaFactory, with full training configurations provided.
When to Use OpenThinker-32B
- Complex Problem Solving: Ideal for applications requiring advanced logical deduction, analytical thinking, and multi-step reasoning.
- Mathematical Applications: Suitable for tasks involving mathematical problem-solving, quantitative analysis, and scientific computing.
- Research & Development: Its fully open-source nature makes it an excellent choice for researchers and developers looking to build upon or analyze reasoning-focused LLMs.
- Benchmarking: Can serve as a strong baseline or comparison model for evaluating reasoning capabilities in new LLM developments.