zurichquants/OpenThinker-7B

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 1, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

OpenThinker-7B by zurichquants is a 7.6 billion parameter language model fine-tuned from Qwen2.5-7B-Instruct. It is specifically optimized for reasoning tasks, leveraging the OpenThoughts-114k dataset, which is distilled from DeepSeek-R1. This model demonstrates improved performance on various reasoning benchmarks, including AIME24 and MATH500, compared to its base model and previous iterations. It is designed for applications requiring enhanced logical and mathematical problem-solving capabilities.

Loading preview...

OpenThinker-7B: Enhanced Reasoning Model

OpenThinker-7B is a 7.6 billion parameter language model developed by zurichquants, built upon the strong foundation of Qwen2.5-7B-Instruct. Its primary distinction lies in its fine-tuning on the extensive OpenThoughts-114k dataset, a high-quality dataset derived from distilling DeepSeek-R1. This specialized training focuses on improving the model's reasoning abilities across various domains.

Key Capabilities & Performance

This model shows notable improvements in reasoning benchmarks, outperforming its predecessor, Bespoke-Stratos-7B. For instance, OpenThinker-7B achieves 31.3 on AIME24 and 83.0 on MATH500, demonstrating enhanced logical and mathematical problem-solving. The development emphasizes full transparency, with all model weights, datasets, and code being open-source. The training involved four 8xH100 nodes over 20 hours, utilizing specific hyperparameters for optimization.

When to Use This Model

  • Reasoning-intensive tasks: Ideal for applications requiring strong logical deduction, mathematical problem-solving, and complex analytical thinking.
  • Research and development: Its open-source nature and detailed training methodology make it suitable for researchers exploring reasoning capabilities in LLMs.
  • Benchmarking: Can be used as a baseline or comparison point for evaluating other models on reasoning-focused datasets like AIME24 and MATH500.