nilgeoutim/RLCR-0.0005smCE-hotpot

TEXT GENERATIONConcurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 15, 2026Architecture:Transformer Featherless Exclusive Cold

The nilgeoutim/RLCR-0.0005smCE-hotpot model is a 3.1 billion parameter language model fine-tuned from Qwen/Qwen2.5-3B. Developed by nilgeoutim, it utilizes the GRPO training method, originally introduced for mathematical reasoning in DeepSeekMath. This model is optimized for enhanced reasoning capabilities, particularly in complex problem-solving scenarios, making it suitable for tasks requiring advanced logical inference.

Loading preview...

Overview

This model, nilgeoutim/RLCR-0.0005smCE-hotpot, is a fine-tuned variant of the Qwen/Qwen2.5-3B base model, featuring 3.1 billion parameters and a 32768-token context length. It was trained using the TRL (Transformer Reinforcement Learning) framework.

Key Capabilities

  • Enhanced Reasoning: The model's training procedure incorporates GRPO (Gradient-based Reward Policy Optimization), a method highlighted in the DeepSeekMath paper for pushing the limits of mathematical reasoning. This suggests an optimization for complex logical and problem-solving tasks.
  • Fine-tuned Performance: As a fine-tuned model, it aims to improve upon the base Qwen2.5-3B's performance in specific areas, likely related to reasoning and structured problem-solving, given the GRPO methodology.

Training Details

The model's training leveraged the TRL framework and the GRPO method, as detailed in the research paper DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. This indicates a focus on reinforcement learning from human feedback or similar reward-based optimization to refine its output quality and reasoning abilities.

Good For

  • Applications requiring advanced logical inference.
  • Tasks benefiting from improved mathematical or complex reasoning.
  • Developers looking for a Qwen2.5-3B derivative with specialized reasoning capabilities.