prithivMLmods/Deepthink-1.5B-Open-PRM

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Apr 23, 2025License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Deepthink-1.5B-Open-PRM is a 1.5 billion parameter process-supervised reasoning model developed by prithivMLmods, fine-tuned from Qwen2.5 1.5B. It specializes in step-by-step mathematical problem-solving in both English and Simplified Chinese, leveraging Process Reward Models (PRM) for interpretable, logically structured responses. This model is optimized for educational applications, STEM tutoring, and lightweight math agents requiring detailed reasoning explanations.

Loading preview...

Deepthink-1.5B-Open-PRM: A Process-Supervised Math Reasoning Model

Deepthink-1.5B-Open-PRM is a 1.5 billion parameter language model developed by prithivMLmods, built upon the Qwen2.5 1.5B architecture. Its core innovation lies in its fine-tuning with Process Reward Models (PRM), which specifically reward high-quality intermediate reasoning steps. This approach fosters step-by-step interpretability and accuracy, making the model's problem-solving process transparent and educational.

Key Capabilities

  • Process Reward Model Supervision: Utilizes PRMs to enhance interpretability and accuracy by focusing on the quality of each reasoning step.
  • Bilingual Math Proficiency: Capable of solving and explaining mathematical problems fluently in both English and Simplified Chinese.
  • Teacher-like Reasoning: Designed to show logical steps before providing an answer, ideal for users who need to understand the 'how' and 'why' of solutions.
  • Long-Context Math Reasoning: Proficient in handling multi-step arithmetic, word problems, logic puzzles, and math up to early college level.
  • Compact and Efficient: Built on Qwen2.5 1.5B, it balances reasoning quality with deployment efficiency, suitable for lightweight environments.

Intended Use Cases

  • Math Education Agents: For creating tutors that provide step-by-step explanations.
  • Bilingual Learning Platforms: Ideal for applications teaching math in both Chinese and English.
  • STEM-Oriented Assistants: Supports early-stage problem-solving in science and engineering.
  • Lightweight LLM Deployments: Optimized for resource-constrained environments like browsers and mobile devices.

Limitations

The model is primarily tuned for math reasoning, meaning performance may degrade on unrelated tasks. Its 1.5B parameter size might struggle with highly abstract or very long multi-domain tasks, and PRM training can introduce biases, necessitating review of outputs.