kikiyaa/Qwen2.5-3B-Instruct-grpo-fullfinetuning-3b-customreward
This model is a 3.1 billion parameter instruction-tuned causal language model, fine-tuned from Qwen/Qwen2.5-3B-Instruct. It was trained using GRPO (Generative Reinforcement Learning with Policy Optimization), a method known for enhancing mathematical reasoning in large language models. This fine-tuning process aims to improve the model's ability to follow instructions and potentially its reasoning capabilities, particularly in areas where GRPO has shown benefits. It is suitable for general instruction-following tasks and applications requiring improved reasoning from a 3B parameter model.
Loading preview...
Model Overview
This model, kikiyaa/Qwen2.5-3B-Instruct-grpo-fullfinetuning-3b-customreward, is a fine-tuned version of the Qwen2.5-3B-Instruct base model, featuring 3.1 billion parameters. It has been specifically trained using GRPO (Generative Reinforcement Learning with Policy Optimization), a method highlighted in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models".
Key Capabilities
- Instruction Following: Enhanced ability to understand and execute user instructions due to its instruction-tuned base and further fine-tuning.
- Reasoning Improvement: Leverages the GRPO training method, which is designed to improve reasoning capabilities, particularly in mathematical contexts as demonstrated by its origin.
- Efficient Deployment: As a 3.1 billion parameter model, it offers a balance between performance and computational efficiency, making it suitable for various applications where larger models might be impractical.
Training Details
The model was fine-tuned using the TRL (Transformer Reinforcement Learning) library. The application of GRPO suggests a focus on refining the model's response generation to align better with desired outcomes, potentially leading to more coherent and logically sound outputs. This approach differentiates it from standard instruction-tuning by incorporating reinforcement learning principles for performance optimization.
Good For
- Applications requiring a compact yet capable instruction-following model.
- Tasks that benefit from improved reasoning, especially if related to structured problem-solving.
- Developers looking for a Qwen2.5-3B variant with a specific GRPO-based fine-tuning approach.