alphadl/R1-Distill-0.6B-Qwen-GRPO

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 3, 2025Architecture:Transformer0.0K Featherless Exclusive Warm

alphadl/R1-Distill-0.6B-Qwen-GRPO is a 0.8 billion parameter language model, fine-tuned from alphadl/R1-Distill-0.6B-Qwen. It was trained using the GRPO method on the OpenR1-Math-220k dataset, specializing it in mathematical reasoning tasks. This model is designed for applications requiring robust mathematical problem-solving capabilities.

Loading preview...

Model Overview

alphadl/R1-Distill-0.6B-Qwen-GRPO is a 0.8 billion parameter language model developed by alphadl. It is a fine-tuned variant of the alphadl/R1-Distill-0.6B-Qwen base model, specifically optimized for mathematical reasoning.

Key Capabilities

  • Mathematical Reasoning: The model has been fine-tuned on the open-r1/OpenR1-Math-220k dataset, enhancing its ability to understand and solve mathematical problems.
  • GRPO Training Method: It leverages the GRPO (Gradient-based Reward Policy Optimization) training method, as introduced in the DeepSeekMath paper, which is designed to push the limits of mathematical reasoning in language models.
  • Efficient Size: With 0.8 billion parameters, it offers a balance between performance in specialized mathematical tasks and computational efficiency.

Good For

  • Mathematical Problem Solving: Ideal for applications requiring the generation of solutions or explanations for mathematical queries.
  • Research in Mathematical LLMs: Useful for researchers exploring the impact of GRPO and specialized mathematical datasets on smaller language models.
  • Educational Tools: Can be integrated into tools for learning or practicing mathematics, providing assistance with problem-solving.