IFM/guru-32B

TEXT GENERATIONPricing:Input $1.06 / Cached $0.053 / Output $2.6Concurrent Unit Cost:2Model Size:32BQuant:FP8Context Size:32kTool Calling:SupportedPublished:May 15, 2025License:cc-by-nc-4.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

IFM/guru-32B is a 32 billion parameter language model based on Qwen2.5-32B, developed by IFM. It is specifically designed and optimized for advanced reasoning across diverse domains including mathematics, coding, science, and logic, achieving high performance on complex benchmarks. With a context length of 32768 tokens, Guru-32B excels in tasks requiring deep analytical capabilities and problem-solving. This model is particularly suited for applications demanding robust reasoning and accurate output in technical and scientific fields.

Loading preview...

IFM/guru-32B: A Reasoning-Optimized Language Model

IFM/guru-32B is a 32 billion parameter model built upon the Qwen2.5-32B architecture, specifically developed for enhanced reasoning capabilities across multiple domains. This model is detailed in the paper "Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective" and demonstrates significant performance improvements in complex analytical tasks.

Key Capabilities and Performance

Guru-32B shows strong performance across a wide array of benchmarks, particularly excelling in:

  • Mathematics: Achieves 34.89 on AIME24 (avg@32) and 86.00 on MATH500, indicating strong mathematical problem-solving skills.
  • Code Generation: Scores 90.85 on HumanEval (avg@4) and 78.80 on MBPP, making it highly proficient in coding tasks.
  • Science & Logic: Demonstrates robust understanding with 50.63 on GPQA-diamond (avg@4) and 45.21 on Zebra Puzzle (avg@4).
  • Tabular Reasoning: Performs well on FinQA (46.14) and HiTab (82.00), showcasing its ability to process and reason over structured data.

Overall, Guru-32B achieves an average score of 54.24 across various reasoning benchmarks, outperforming other models like ORZ 32B and SimpleRL 32B in many categories. The model's evaluation methodology uses specific parameters (temperature=1.0, top_p=0.7) for consistent comparison.

Ideal Use Cases

This model is particularly well-suited for applications requiring:

  • Advanced Problem Solving: In fields like engineering, scientific research, and quantitative analysis.
  • High-Accuracy Code Generation: For developers needing reliable and complex code solutions.
  • Complex Data Interpretation: Especially for tasks involving tabular data or intricate logical puzzles.
  • Educational Tools: For generating explanations or solving problems in STEM subjects.