nilgeoutim/RLCR-0.00005smCE-hotpot-seed40

TEXT GENERATIONConcurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 31, 2026Architecture:Transformer Featherless Exclusive Cold

The nilgeoutim/RLCR-0.00005smCE-hotpot-seed40 is a 3.1 billion parameter language model, fine-tuned from Qwen/Qwen2.5-3B. It was trained using the TRL framework and incorporates the GRPO method, which is designed to enhance mathematical reasoning. This model is optimized for tasks requiring advanced reasoning capabilities, particularly in mathematical contexts, leveraging its specialized training approach.

Loading preview...

Model Overview

nilgeoutim/RLCR-0.00005smCE-hotpot-seed40 is a 3.1 billion parameter language model built upon the Qwen/Qwen2.5-3B architecture. It has been fine-tuned using the TRL (Transformer Reinforcement Learning) framework, specifically incorporating the GRPO (Gradient-based Reinforcement Learning with Policy Optimization) method. This training approach is derived from research presented in "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300).

Key Capabilities

  • Enhanced Mathematical Reasoning: The primary differentiator of this model is its specialized training with GRPO, aiming to improve performance on tasks requiring complex mathematical reasoning.
  • Qwen2.5-3B Base: Benefits from the foundational capabilities of the Qwen2.5-3B model, providing a strong base for general language understanding and generation.
  • TRL Framework: Utilizes the TRL library for efficient and effective fine-tuning processes.

Training Details

The model's training procedure involved the GRPO method, as detailed in the DeepSeekMath paper. The development environment included TRL 0.16.0.dev0, Transformers 4.48.3, Pytorch 2.5.1+cu124, Datasets 4.0.0, and Tokenizers 0.21.1.

Good For

  • Applications requiring improved mathematical problem-solving and reasoning.
  • Research and development in reinforcement learning from human feedback (RLHF) applied to language models.
  • Tasks where a 3.1 billion parameter model with a 32768 token context length is suitable for balancing performance and computational resources.