roaminwind/DeepSeek-R1-Distill-Qwen-1.5B-GRPO

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Apr 13, 2025Architecture:Transformer Featherless Exclusive Cold

roaminwind/DeepSeek-R1-Distill-Qwen-1.5B-GRPO is a 1.5 billion parameter language model fine-tuned from deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B. It was trained using the GRPO method on the OpenR1-Math-220k dataset, specializing it for mathematical reasoning tasks. This model leverages techniques from DeepSeekMath to enhance its capabilities in complex mathematical problem-solving. Its primary strength lies in mathematical reasoning, making it suitable for applications requiring robust numerical and logical processing.

Loading preview...

Model Overview

roaminwind/DeepSeek-R1-Distill-Qwen-1.5B-GRPO is a specialized language model with 1.5 billion parameters, built upon the deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B architecture. This model has undergone further fine-tuning specifically for mathematical reasoning tasks.

Key Differentiators

  • Mathematical Reasoning Focus: The model was fine-tuned on the open-r1/OpenR1-Math-220k dataset, which is dedicated to mathematical problems.
  • GRPO Training Method: It utilizes the GRPO (Guided Reinforcement Learning with Policy Optimization) training method, as introduced in the DeepSeekMath paper. This method is designed to push the limits of mathematical reasoning in language models.
  • TRL Framework: Training was conducted using the Hugging Face TRL library, a framework for Transformer Reinforcement Learning.

Intended Use Cases

This model is particularly well-suited for applications requiring:

  • Solving mathematical problems.
  • Generating logical steps for mathematical solutions.
  • Tasks that benefit from enhanced numerical and logical reasoning capabilities.