nilgeoutim/RLCR-0.025smCE-hotpot

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 11, 2026Architecture:Transformer Featherless Exclusive Cold

The nilgeoutim/RLCR-0.025smCE-hotpot model is a 3.1 billion parameter language model, fine-tuned from Qwen/Qwen2.5-3B using the TRL framework. It was trained with GRPO, a method specifically designed to enhance mathematical reasoning capabilities, as introduced in the DeepSeekMath paper. This model is optimized for complex reasoning tasks, particularly those involving mathematical problem-solving, and offers a 32768 token context length. Its primary strength lies in its ability to process and generate responses for intricate logical and mathematical queries.

Loading preview...

Model Overview

The nilgeoutim/RLCR-0.025smCE-hotpot is a 3.1 billion parameter language model, fine-tuned from the Qwen/Qwen2.5-3B base model. It leverages the TRL (Transformer Reinforcement Learning) framework for its training process.

Key Capabilities

  • Enhanced Mathematical Reasoning: This model was specifically trained using GRPO (Gradient-based Reinforcement Learning with Policy Optimization), a method detailed in the DeepSeekMath paper. This training approach aims to significantly improve its ability to handle and solve complex mathematical reasoning problems.
  • Extended Context Length: It supports a substantial context window of 32768 tokens, allowing for the processing of longer and more detailed inputs.
  • Instruction Following: As a fine-tuned model, it is designed to follow instructions effectively, making it suitable for various conversational and question-answering tasks.

Training Details

The model's training procedure utilized GRPO, a technique focused on pushing the limits of mathematical reasoning in open language models. The training environment included:

  • TRL: 0.16.0.dev0
  • Transformers: 4.48.3
  • Pytorch: 2.5.1+cu124
  • Datasets: 4.0.0
  • Tokenizers: 0.21.1

Good for

  • Mathematical Problem Solving: Ideal for applications requiring advanced mathematical reasoning and problem-solving.
  • Complex Reasoning Tasks: Suitable for scenarios where logical deduction and intricate thought processes are needed.
  • Research and Development: Can serve as a base for further experimentation and fine-tuning on specific reasoning-intensive datasets.