logan7000/cogrpo-n3-ring-qwen25-3b-x-llama32-3b-x-granite2b-math345-groupB-llama32-end

TEXT GENERATIONPricing:Input $0.2036 / Output $1.34Concurrent Unit Cost:1Model Size:3.2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 27, 2026Architecture:Transformer Featherless Exclusive Cold

This model, named group_B, is a 3.2 billion parameter language model fine-tuned from meta-llama/Llama-3.2-3B-Instruct. It was trained using the TRL framework and incorporates the GRPO method, which is designed to enhance mathematical reasoning capabilities. This model is specifically optimized for tasks requiring advanced mathematical problem-solving and logical deduction.

Loading preview...

Overview

This model, designated as group_B, is a 3.2 billion parameter language model that has been fine-tuned from the meta-llama/Llama-3.2-3B-Instruct base model. The fine-tuning process utilized the TRL (Transformers Reinforcement Learning) framework.

Key Capabilities

  • Enhanced Mathematical Reasoning: A core differentiator of this model is its training with the GRPO (Gradient-based Reinforcement Learning for Policy Optimization) method. This technique, introduced in the DeepSeekMath paper, is specifically aimed at improving the model's ability to handle complex mathematical reasoning tasks.
  • Instruction Following: As it is fine-tuned from an instruct model, it is designed to follow user instructions effectively.

Training Details

The model's training procedure leveraged the GRPO method, as detailed in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). The training environment included specific versions of key frameworks:

  • TRL: 1.2.0.dev0
  • Transformers: 4.57.6
  • Pytorch: 2.10.0+cu128
  • Datasets: 5.0.1
  • Tokenizers: 0.22.2

Good For

  • Applications requiring strong mathematical problem-solving.
  • Tasks that benefit from advanced logical reasoning.
  • Developers looking for a compact yet capable model for instruction-tuned mathematical and reasoning-based queries.