nilgeoutim/RLCR-smCE-hotpot

TEXT GENERATIONConcurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 7, 2026Architecture:Transformer Featherless Exclusive Cold

nilgeoutim/RLCR-smCE-hotpot is a 3.1 billion parameter causal language model, fine-tuned from Qwen/Qwen2.5-3B. This model was trained using the GRPO method, as introduced in the DeepSeekMath paper, which focuses on enhancing mathematical reasoning capabilities. It is optimized for complex reasoning tasks, particularly those benefiting from advanced reinforcement learning techniques. The model leverages a 32768 token context length to process extensive inputs for detailed problem-solving.

Loading preview...

Model Overview

nilgeoutim/RLCR-smCE-hotpot is a 3.1 billion parameter language model, fine-tuned from the Qwen/Qwen2.5-3B base model. It utilizes a substantial 32768 token context window, allowing it to process and generate longer, more complex sequences of text. The model's training incorporated the TRL framework and specifically employed the GRPO (Generalized Reinforcement Learning with Policy Optimization) method.

Key Capabilities

  • Enhanced Reasoning: Fine-tuned with the GRPO method, which is designed to improve mathematical and general reasoning abilities, as detailed in the DeepSeekMath research paper.
  • Large Context Window: Benefits from a 32768 token context length, suitable for tasks requiring extensive input understanding or generating detailed responses.
  • Qwen2.5 Architecture: Built upon the robust Qwen2.5-3B architecture, providing a strong foundation for language understanding and generation.

Good For

  • Complex Problem Solving: Ideal for applications that demand advanced reasoning, potentially including mathematical or logical challenges.
  • Detailed Text Generation: Its large context window makes it suitable for generating comprehensive and coherent long-form content.
  • Research and Development: Useful for researchers exploring the impact of GRPO and similar reinforcement learning techniques on LLM performance.