kikiyaa/Qwen2.5-3B-Instruct-grpo-fullfinetuning-3b-customreward

TEXT GENERATIONConcurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 7, 2026Architecture:Transformer Featherless Exclusive Cold

This model is a 3.1 billion parameter instruction-tuned causal language model, fine-tuned from Qwen/Qwen2.5-3B-Instruct. It was trained using GRPO (Generative Reinforcement Learning with Policy Optimization), a method known for enhancing mathematical reasoning in large language models. This fine-tuning process aims to improve the model's ability to follow instructions and potentially its reasoning capabilities, particularly in areas where GRPO has shown benefits. It is suitable for general instruction-following tasks and applications requiring improved reasoning from a 3B parameter model.

Loading preview...

Model Overview

This model, kikiyaa/Qwen2.5-3B-Instruct-grpo-fullfinetuning-3b-customreward, is a fine-tuned version of the Qwen2.5-3B-Instruct base model, featuring 3.1 billion parameters. It has been specifically trained using GRPO (Generative Reinforcement Learning with Policy Optimization), a method highlighted in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models".

Key Capabilities

  • Instruction Following: Enhanced ability to understand and execute user instructions due to its instruction-tuned base and further fine-tuning.
  • Reasoning Improvement: Leverages the GRPO training method, which is designed to improve reasoning capabilities, particularly in mathematical contexts as demonstrated by its origin.
  • Efficient Deployment: As a 3.1 billion parameter model, it offers a balance between performance and computational efficiency, making it suitable for various applications where larger models might be impractical.

Training Details

The model was fine-tuned using the TRL (Transformer Reinforcement Learning) library. The application of GRPO suggests a focus on refining the model's response generation to align better with desired outcomes, potentially leading to more coherent and logically sound outputs. This approach differentiates it from standard instruction-tuning by incorporating reinforcement learning principles for performance optimization.

Good For

  • Applications requiring a compact yet capable instruction-following model.
  • Tasks that benefit from improved reasoning, especially if related to structured problem-solving.
  • Developers looking for a Qwen2.5-3B variant with a specific GRPO-based fine-tuning approach.