shaffhausenquant/OpenThinker-7B
OpenThinker-7B by shaffhausenquant is a 7.6 billion parameter language model, fine-tuned from Qwen/Qwen2.5-7B-Instruct on the OpenThoughts-114k dataset. This model is specifically optimized for reasoning tasks, demonstrating improved performance over its base model and Bespoke-Stratos-7B on benchmarks like AIME24, MATH500, and GPQA-Diamond. It is designed for applications requiring advanced problem-solving and analytical capabilities.
Loading preview...
OpenThinker-7B: Enhanced Reasoning Model
OpenThinker-7B is a 7.6 billion parameter language model developed by shaffhausenquant, built upon the robust Qwen2.5-7B-Instruct architecture. Its primary distinction lies in its fine-tuning on the extensive OpenThoughts-114k dataset, which was distilled from DeepSeek-R1, focusing on complex reasoning examples.
Key Capabilities & Performance
This model demonstrates significant improvements in reasoning benchmarks compared to its predecessor, Bespoke-Stratos-7B, and the base Qwen model. Key performance highlights include:
- AIME24: Achieves 31.3, outperforming Bespoke-Stratos-7B (22.7).
- MATH500: Scores 83.0, an improvement over Bespoke-Stratos-7B (79.6).
- GPQA-Diamond: Reaches 42.4, surpassing Bespoke-Stratos-7B (38.9).
- LCBv2: Shows enhanced performance across easy, medium, and hard categories, with an overall score of 39.9.
Training and Open-Source Commitment
OpenThinker-7B was trained for 20 hours using four 8xH100 nodes, utilizing specific hyperparameters including a learning rate of 1e-05 and a total batch size of 96. The project emphasizes full transparency and open-source availability, with its model weights, datasets, data generation code, evaluation code (Evalchemy), and training code all publicly accessible. This commitment allows for community inspection and further development.
Good for:
- Advanced Reasoning Tasks: Excels in mathematical problem-solving, scientific reasoning, and complex question answering.
- Research and Development: Provides a strong open-source foundation for exploring and building upon reasoning-focused LLMs.
- Benchmarking: Useful for evaluating and comparing performance on challenging reasoning datasets.