nilgeoutim/RLCR-0.025smCE-hotpot
The nilgeoutim/RLCR-0.025smCE-hotpot model is a 3.1 billion parameter language model, fine-tuned from Qwen/Qwen2.5-3B using the TRL framework. It was trained with GRPO, a method specifically designed to enhance mathematical reasoning capabilities, as introduced in the DeepSeekMath paper. This model is optimized for complex reasoning tasks, particularly those involving mathematical problem-solving, and offers a 32768 token context length. Its primary strength lies in its ability to process and generate responses for intricate logical and mathematical queries.
Loading preview...
Model Overview
The nilgeoutim/RLCR-0.025smCE-hotpot is a 3.1 billion parameter language model, fine-tuned from the Qwen/Qwen2.5-3B base model. It leverages the TRL (Transformer Reinforcement Learning) framework for its training process.
Key Capabilities
- Enhanced Mathematical Reasoning: This model was specifically trained using GRPO (Gradient-based Reinforcement Learning with Policy Optimization), a method detailed in the DeepSeekMath paper. This training approach aims to significantly improve its ability to handle and solve complex mathematical reasoning problems.
- Extended Context Length: It supports a substantial context window of 32768 tokens, allowing for the processing of longer and more detailed inputs.
- Instruction Following: As a fine-tuned model, it is designed to follow instructions effectively, making it suitable for various conversational and question-answering tasks.
Training Details
The model's training procedure utilized GRPO, a technique focused on pushing the limits of mathematical reasoning in open language models. The training environment included:
- TRL: 0.16.0.dev0
- Transformers: 4.48.3
- Pytorch: 2.5.1+cu124
- Datasets: 4.0.0
- Tokenizers: 0.21.1
Good for
- Mathematical Problem Solving: Ideal for applications requiring advanced mathematical reasoning and problem-solving.
- Complex Reasoning Tasks: Suitable for scenarios where logical deduction and intricate thought processes are needed.
- Research and Development: Can serve as a base for further experimentation and fine-tuning on specific reasoning-intensive datasets.