logan7000/llm-math345-ttrl-granite2b-endpoint

TEXT GENERATIONPricing:Input $0.32 / Cached $0.016 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kPublished:Aug 25, 2026Architecture:Transformer Featherless Exclusive Cold

This model is a 2 billion parameter instruction-tuned language model, fine-tuned from ibm-granite/granite-3.3-2b-instruct. It was trained using the GRPO method, which is designed to enhance mathematical reasoning in language models. This makes it particularly suitable for tasks requiring robust mathematical and logical problem-solving capabilities. The model leverages the TRL framework for its training procedure.

Loading preview...

Model Overview

This model is a fine-tuned version of the ibm-granite/granite-3.3-2b-instruct model, featuring 2 billion parameters. Its training incorporated the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). The training process utilized the TRL (Transformers Reinforcement Learning) framework.

Key Characteristics

  • Base Model: ibm-granite/granite-3.3-2b-instruct
  • Parameter Count: 2 billion
  • Training Method: GRPO, focused on improving mathematical reasoning.
  • Framework: TRL (Transformers Reinforcement Learning).

Ideal Use Cases

This model is well-suited for applications that require:

  • Mathematical Reasoning: Tasks involving numerical problems, logical deductions, and mathematical problem-solving.
  • Instruction Following: Responding to user prompts and instructions effectively, building upon its instruction-tuned base.
  • Research and Development: Exploring the capabilities of models trained with advanced reinforcement learning techniques like GRPO for specific domains.