logan7000/cogrpo-n3-ring-qwen25-3b-x-llama32-3b-x-granite2b-math345-groupC-granite2b-end

TEXT GENERATIONPricing:Input $0.32 / Cached $0.016 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kPublished:Aug 27, 2026Architecture:Transformer Featherless Exclusive Cold

This model is a 2 billion parameter language model, fine-tuned from ibm-granite/granite-3.3-2b-instruct. It was trained using the GRPO method, which is designed to enhance mathematical reasoning capabilities. This model is particularly suited for tasks requiring advanced mathematical problem-solving and logical deduction.

Loading preview...

Model Overview

This model is a fine-tuned version of the ibm-granite/granite-3.3-2b-instruct base model, featuring 2 billion parameters and a 32768 token context length. It has been specifically trained using the GRPO (Gradient Regularized Policy Optimization) method, as introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models". This training approach aims to significantly improve the model's performance in mathematical reasoning tasks.

Key Capabilities

  • Enhanced Mathematical Reasoning: Optimized through the GRPO method for complex mathematical problem-solving.
  • Instruction Following: Built upon an instruct-tuned base model, it is designed to follow user instructions effectively.
  • Fine-tuned Performance: Leverages the TRL (Transformers Reinforcement Learning) framework for its fine-tuning process.

When to Use This Model

This model is ideal for applications requiring strong mathematical and logical reasoning. Consider using it for:

  • Solving mathematical problems.
  • Tasks that benefit from advanced logical deduction.
  • Research and development in improving LLM mathematical capabilities.