logan7000/cogrpo-n3-ring-qwen25-3b-x-llama32-3b-x-phi4mini-math345-groupA-qwen25-end

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 25, 2026Architecture:Transformer Featherless Exclusive Cold

This model is a 3.1 billion parameter language model fine-tuned from Qwen/Qwen2.5-3B. It was trained using the GRPO method, which is designed to enhance mathematical reasoning capabilities. This makes it particularly suitable for tasks requiring advanced mathematical problem-solving and logical deduction.

Loading preview...

Model Overview

This model, named group_A, is a fine-tuned version of the Qwen/Qwen2.5-3B language model, featuring approximately 3.1 billion parameters. It has been specifically trained using the GRPO (Grouped Reinforcement Learning with Policy Optimization) method, as introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models".

Key Capabilities

  • Enhanced Mathematical Reasoning: The application of the GRPO training procedure suggests a focus on improving the model's ability to handle complex mathematical problems and logical reasoning tasks.
  • Instruction Following: As a fine-tuned model, it is designed to follow instructions effectively, making it suitable for various prompt-based applications.

Training Details

The model's training leveraged the TRL (Transformers Reinforcement Learning) library. The GRPO method, central to its training, aims to optimize performance in areas like mathematical problem-solving. This approach differentiates it from standard instruction-tuned models by emphasizing a specific reasoning paradigm.

Good For

  • Applications requiring strong mathematical reasoning.
  • Tasks involving complex logical deduction.
  • Research and development in advanced language model training techniques, particularly GRPO.