nilgeoutim/RLCR-0.005smCE-hotpot

TEXT GENERATIONConcurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 14, 2026Architecture:Transformer Featherless Exclusive Cold

The nilgeoutim/RLCR-0.005smCE-hotpot model is a 3.1 billion parameter language model, fine-tuned from Qwen/Qwen2.5-3B. It was trained using the GRPO method, as introduced in the DeepSeekMath paper, which focuses on enhancing mathematical reasoning. This model is optimized for tasks requiring advanced reasoning capabilities, leveraging its 32768-token context length. Its primary strength lies in complex problem-solving, particularly in areas benefiting from structured reasoning.

Loading preview...

Overview

nilgeoutim/RLCR-0.005smCE-hotpot is a 3.1 billion parameter language model, building upon the Qwen/Qwen2.5-3B architecture. It has been specifically fine-tuned using the TRL framework and incorporates the GRPO (Gradient-based Reward Policy Optimization) training method. GRPO, detailed in the "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" paper, is designed to significantly improve a model's mathematical and general reasoning abilities.

Key Capabilities

  • Enhanced Reasoning: Leverages the GRPO training method to improve performance on tasks requiring logical and mathematical reasoning.
  • Qwen2.5-3B Base: Benefits from the robust foundational capabilities of the Qwen2.5-3B model.
  • Extended Context: Supports a substantial context length of 32768 tokens, allowing for processing and understanding longer inputs and complex problem descriptions.

Good For

  • Mathematical Problem Solving: Ideal for applications that involve complex mathematical equations, proofs, or quantitative analysis.
  • Logical Reasoning Tasks: Suitable for scenarios requiring structured thought processes and deductive reasoning.
  • Research and Development: Can be used as a base for further experimentation and fine-tuning on specific reasoning-intensive datasets.