zbeeb/Qwen2.5-3B-OpenR1-SFT
The zbeeb/Qwen2.5-3B-OpenR1-SFT is a 3.1 billion parameter language model built upon the Qwen2.5-3B architecture, fine-tuned specifically for mathematical reasoning. This model is the result of a supervised fine-tuning (SFT) stage from a token-level study, focusing on the transition from SFT to GRPO. It was trained on a selected subset of the OpenR1-Math-220k dataset, making it particularly adept at solving mathematical problems and generating step-by-step reasoning. Its primary use case is in applications requiring robust mathematical problem-solving capabilities.
Loading preview...
Model Overview
zbeeb/Qwen2.5-3B-OpenR1-SFT is a 3.1 billion parameter language model derived from the Qwen2.5-3B base model. It represents the final supervised fine-tuning (SFT) checkpoint from a study investigating the transition from SFT to GRPO (Generative Reinforcement Learning with Policy Optimization). The model's weights were specifically modified through SFT to enhance its mathematical reasoning abilities.
Key Capabilities
- Mathematical Reasoning: Fine-tuned on a subset of the OpenR1-Math-220k dataset, the model excels at processing and generating mathematical solutions.
- Step-by-Step Problem Solving: Designed to provide detailed, step-by-step reasoning for mathematical problems, with prompt tokens masked from the supervised loss during training.
- Optimized Training: The SFT stage involved 278 optimizer updates and 20,016 examples, utilizing a 4096-token sequence length and
bf16precision.
Use Cases
This model is particularly well-suited for applications requiring:
- Automated mathematical problem solving.
- Educational tools that need to explain mathematical concepts or solutions.
- Research into improving language models' logical and mathematical reasoning capabilities.
Limitations
It's important to note that the provided evaluation-summary.jsonl contains experiment probe results and does not claim standard benchmark performance. The model's primary focus is on mathematical reasoning, and its performance on other general language tasks may not be optimized.