nilgeoutim/RLCR-0.25smCE-hotpot

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 12, 2026Architecture:Transformer Featherless Exclusive Cold

nilgeoutim/RLCR-0.25smCE-hotpot is a 3.1 billion parameter language model fine-tuned from Qwen/Qwen2.5-3B. Developed by nilgeoutim, this model utilizes the GRPO method, known for enhancing mathematical reasoning in large language models. It is specifically optimized for tasks requiring advanced reasoning, building upon its base architecture with a 32768 token context length.

Loading preview...

Model Overview

nilgeoutim/RLCR-0.25smCE-hotpot is a 3.1 billion parameter language model, fine-tuned from the Qwen/Qwen2.5-3B base model. This model was developed by nilgeoutim using the TRL library and incorporates the GRPO (Gradient-based Reinforcement Learning with Policy Optimization) method. GRPO is a technique introduced in the context of DeepSeekMath, aiming to push the limits of mathematical reasoning in open language models.

Key Capabilities

  • Enhanced Reasoning: Leverages the GRPO training method to improve reasoning capabilities, particularly in areas where mathematical understanding is beneficial.
  • Qwen2.5-3B Foundation: Builds upon the robust architecture and pre-training of the Qwen2.5-3B model, providing a strong general language understanding base.
  • Extended Context Length: Supports a context length of 32768 tokens, allowing for processing and generating longer sequences of text.

Good For

  • Reasoning-intensive tasks: Suitable for applications requiring logical deduction or problem-solving, potentially benefiting from the GRPO-enhanced training.
  • General text generation: Can be used for a wide range of natural language processing tasks due to its Qwen2.5-3B lineage.
  • Exploration of GRPO: Provides an accessible model for developers interested in experimenting with models trained using the GRPO methodology for improved reasoning.