luckeciano/Qwen-2.5-7B-RL-LACPO-NoBaselineAlpha1
The luckeciano/Qwen-2.5-7B-RL-LACPO-NoBaselineAlpha1 is a 7.6 billion parameter Qwen 2.5 model, fine-tuned from Qwen/Qwen2.5-Math-7B. It was trained using the GRPO method, as introduced in the DeepSeekMath paper, on the DigitalLearningGmbH/MATH-lighteval dataset. This model is specifically optimized for advanced mathematical reasoning tasks, leveraging its 32768-token context length.
Loading preview...
Model Overview
This model, luckeciano/Qwen-2.5-7B-RL-LACPO-NoBaselineAlpha1, is a 7.6 billion parameter language model built upon the Qwen 2.5 architecture. It is a fine-tuned variant of the Qwen/Qwen2.5-Math-7B base model, specifically enhanced for mathematical reasoning capabilities.
Key Training Details
- Base Model: Qwen/Qwen2.5-Math-7B
- Fine-tuning Dataset: DigitalLearningGmbH/MATH-lighteval
- Training Method: Utilizes GRPO (Generalized Reinforcement Learning with Policy Optimization), a technique detailed in the DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models paper.
- Framework: Trained using the TRL library (Transformer Reinforcement Learning).
Primary Use Case
This model is particularly well-suited for applications requiring strong mathematical problem-solving and reasoning. Its fine-tuning on a dedicated math dataset and the application of the GRPO method suggest improved performance in generating accurate and logical responses to complex mathematical queries.