logan7000/llm-math345-gt-phi4mini-endpoint

TEXT GENERATIONPricing:Input $0.32 / Cached $0.016 / Output $1.4Concurrent Unit Cost:1Model Size:3.8BQuant:BF16Context Size:32kPublished:Aug 25, 2026Architecture:Transformer Featherless Exclusive Cold

This model, developed by gt, is a 3.8 billion parameter instruction-tuned variant of microsoft/Phi-4-mini-instruct, featuring a 32768-token context length. It was fine-tuned using the TRL framework and the GRPO method, specifically optimized for mathematical reasoning tasks. This specialization makes it particularly effective for complex numerical and logical problem-solving.

Loading preview...

Model Overview

This model is a fine-tuned version of the microsoft/Phi-4-mini-instruct base model, developed by gt. It leverages a 3.8 billion parameter architecture and supports a substantial context length of 32768 tokens, making it suitable for processing longer inputs.

Key Capabilities

  • Mathematical Reasoning: The model was specifically trained using the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models." This training methodology enhances its ability to handle complex mathematical and logical problems.
  • Instruction Following: As an instruction-tuned model, it is designed to accurately interpret and respond to user prompts and instructions.
  • TRL Framework: Fine-tuned with the TRL (Transformers Reinforcement Learning) library, indicating a focus on optimizing model behavior through reinforcement learning techniques.

When to Use This Model

  • Mathematical Problem Solving: Ideal for applications requiring strong mathematical reasoning, such as solving equations, logical puzzles, or quantitative analysis.
  • Research and Development: Useful for researchers exploring the application of GRPO and reinforcement learning in enhancing LLM capabilities for specific domains.
  • Instruction-based Tasks: Suitable for general instruction-following tasks where the underlying mathematical reasoning capabilities can provide robust and accurate responses.