Blancy/Qwen-2.5-7B-Simple-RL

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Feb 26, 2025Architecture:Transformer0.0K Featherless Exclusive Cold

Blancy/Qwen-2.5-7B-Simple-RL is a 7.6 billion parameter language model developed by Blancy, fine-tuned from Qwen/Qwen2.5-Math-7B. This model leverages the GRPO training method, known for enhancing mathematical reasoning in large language models. With a context length of 32768 tokens, it is specifically optimized for tasks requiring advanced mathematical problem-solving capabilities.

Loading preview...

Model Overview

Blancy/Qwen-2.5-7B-Simple-RL is a 7.6 billion parameter language model, fine-tuned from the Qwen/Qwen2.5-Math-7B base model. It was developed by Blancy using the TRL framework.

Key Differentiator: GRPO Training

This model's primary distinction lies in its training methodology. It was trained using GRPO (Gradient-based Reward Policy Optimization), a method introduced in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models". This technique is specifically designed to enhance and push the boundaries of mathematical reasoning capabilities in large language models.

Training Details

The fine-tuning process utilized the Blancy/secondfiltered-math220k-difficulty_stratified_10k dataset. The training was conducted using the TRL library, with specific framework versions including TRL 0.15.2, Transformers 4.49.0, Pytorch 2.5.1, Datasets 3.5.0, and Tokenizers 0.21.1.

Use Cases

Given its specialized training with GRPO and a math-focused dataset, this model is particularly well-suited for applications requiring:

  • Advanced mathematical problem-solving
  • Reasoning tasks involving numerical and logical operations
  • Educational tools for mathematics

Its 32768-token context length further supports handling complex and lengthy mathematical problems or discussions.