nilgeoutim/RLCR-0.00005smCE-hotpot-seed44
TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 5, 2026Architecture:Transformer Featherless Exclusive Cold
nilgeoutim/RLCR-0.00005smCE-hotpot-seed44 is a 3.1 billion parameter language model fine-tuned from Qwen/Qwen2.5-3B. Developed by nilgeoutim, this model was trained using the GRPO method, which is designed to enhance mathematical reasoning capabilities. It is optimized for tasks requiring advanced reasoning, leveraging techniques from DeepSeekMath research. This model is suitable for applications demanding improved logical and mathematical problem-solving from a compact language model.
Loading preview...
Model Overview
nilgeoutim/RLCR-0.00005smCE-hotpot-seed44 is a 3.1 billion parameter language model, fine-tuned from the Qwen/Qwen2.5-3B base model. This model was developed by nilgeoutim and utilizes the TRL framework for its training process.
Key Training Details
- Fine-tuning Method: The model was trained using GRPO (Gradient-based Reward Policy Optimization), a method introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models". This suggests a focus on enhancing the model's capabilities in mathematical and logical reasoning tasks.
- Frameworks: Training was conducted with TRL (Transformer Reinforcement Learning) version 0.16.0.dev0, alongside Transformers 4.48.3 and Pytorch 2.5.1+cu124.
Potential Use Cases
- Mathematical Reasoning: Given its training with the GRPO method, the model is likely well-suited for tasks that require mathematical problem-solving, logical deduction, and complex reasoning.
- Question Answering: Its base in Qwen2.5-3B combined with specialized training could make it effective for detailed question answering, particularly in domains benefiting from enhanced reasoning.