Hao0oWang/Qwen2.5-Math-7B-16k-think

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Dec 28, 2025License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

Hao0oWang/Qwen2.5-Math-7B-16k-think is a 7.6 billion parameter language model based on Qwen2.5-Math-7B-Base, developed by Hao Wang et al. It is specifically fine-tuned for long context reasoning, particularly in mathematical tasks, and features an extended context window of 32768 tokens. This model incorporates modifications to its chat template and rope_theta for enhanced performance in complex reasoning scenarios. Its primary application is in advanced mathematical problem-solving and reasoning tasks requiring extensive contextual understanding.

Loading preview...

Model Overview

Hao0oWang/Qwen2.5-Math-7B-16k-think is a 7.6 billion parameter model derived from the Qwen2.5-Math-7B-Base, developed by Hao Wang and his team. This model is a key component of the CurioSFT framework, which focuses on entropy-preserving supervised fine-tuning for large reasoning models, as detailed in their research paper, "Learning While Staying Curious: Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models."

Key Enhancements

This model distinguishes itself through several critical modifications aimed at improving its reasoning capabilities, especially for long-context scenarios:

  • Extended Context Window: The model's context window has been significantly expanded to 32768 tokens, enabling it to process and understand much longer inputs.
  • Optimized for Reasoning: Specific adjustments to the chat_template and rope_theta have been implemented to enhance its performance in complex reasoning tasks.
  • Mathematical Focus: Built upon a math-specific base model, it is particularly adept at handling mathematical problems and logical deductions.

Ideal Use Cases

This model is particularly well-suited for applications requiring:

  • Advanced Mathematical Problem Solving: Excelling in tasks that demand deep mathematical understanding and computation.
  • Long-Context Reasoning: Effectively processing and reasoning over extensive textual information, beneficial for complex analytical tasks.
  • Research in Reasoning Models: Serving as a foundational model for further research and development in large reasoning models, especially within the context of adaptive self-distillation techniques.