swadeshb/Llama-3.2-3B-Instruct-TRACE_GRPO-Final

TEXT GENERATIONPricing:Input $0.2036 / Output $1.34Concurrent Unit Cost:1Model Size:3.2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Nov 13, 2025Architecture:Transformer Featherless Exclusive Cold

swadeshb/Llama-3.2-3B-Instruct-TRACE_GRPO-Final is a 3.2 billion parameter instruction-tuned causal language model, fine-tuned from meta-llama/Llama-3.2-3B-Instruct. This model was trained using the GRPO method, which is designed to enhance mathematical reasoning capabilities. It is optimized for tasks requiring improved mathematical problem-solving and logical deduction, making it suitable for applications in technical and scientific domains.

Loading preview...

Model Overview

This model, swadeshb/Llama-3.2-3B-Instruct-TRACE_GRPO-Final, is a specialized instruction-tuned variant of the meta-llama/Llama-3.2-3B-Instruct base model. It features 3.2 billion parameters and has been fine-tuned using the TRL framework.

Key Capabilities

  • Enhanced Mathematical Reasoning: The model's primary differentiator is its training with the GRPO (Gradient-based Reasoning Policy Optimization) method. This technique, introduced in the DeepSeekMath paper, aims to significantly improve the model's ability to handle mathematical reasoning tasks.
  • Instruction Following: As an instruction-tuned model, it is designed to understand and execute user prompts effectively, providing relevant and coherent responses.

When to Use This Model

  • Mathematical Problem Solving: Ideal for applications requiring robust mathematical reasoning, such as solving equations, logical deductions, or generating explanations for mathematical concepts.
  • Research and Development: Suitable for researchers exploring advanced fine-tuning techniques for improving specific cognitive abilities in LLMs, particularly in the domain of mathematics.
  • Specialized Instruction-Following: Use cases where a smaller, yet capable, instruction-tuned model with a focus on analytical tasks is preferred.