dekangli/Qwen2.5-1.5B-GRPO-v5

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:May 24, 2025Architecture:Transformer Featherless Exclusive Cold

dekangli/Qwen2.5-1.5B-GRPO-v5 is a 1.5 billion parameter language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It was trained using the GRPO method on the OpenR1-Math-220k dataset, specializing it for mathematical reasoning tasks. This model is optimized to enhance performance in complex mathematical problem-solving.

Loading preview...

Model Overview

This model, dekangli/Qwen2.5-1.5B-GRPO-v5, is a specialized 1.5 billion parameter language model. It is built upon the Qwen/Qwen2.5-1.5B-Instruct architecture and has been further fine-tuned to excel in specific domains.

Key Capabilities

  • Mathematical Reasoning: The model's primary strength lies in mathematical problem-solving, achieved through fine-tuning on the open-r1/OpenR1-Math-220k dataset.
  • GRPO Training: It leverages the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the DeepSeekMath paper, to enhance its reasoning abilities.

Training Details

The model was trained using the TRL framework (version 0.18.0.dev0) with Transformers 4.52.0.dev0 and Pytorch 2.6.0. The GRPO method is detailed in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300).

When to Use This Model

This model is particularly well-suited for applications requiring robust mathematical reasoning and problem-solving, especially where the Qwen2.5-1.5B-Instruct base model needs enhanced mathematical capabilities.