sergiopaniego/qwen3-1.7b-mbpp-grpo
TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 19, 2026Architecture:Transformer Featherless Exclusive Cold
The sergiopaniego/qwen3-1.7b-mbpp-grpo model is a 1.7 billion parameter language model, fine-tuned from Qwen/Qwen3-1.7B by sergiopaniego. It was trained using the GRPO method, which is designed to enhance mathematical reasoning capabilities in language models. This model is particularly suited for tasks requiring improved logical and mathematical problem-solving, leveraging its specialized training approach.
Loading preview...
Model Overview
sergiopaniego/qwen3-1.7b-mbpp-grpo is a 1.7 billion parameter language model, fine-tuned by sergiopaniego from the base Qwen/Qwen3-1.7B architecture. This model incorporates a specialized training methodology known as GRPO (Gradient-based Reward Policy Optimization).
Key Characteristics
- Base Model: Fine-tuned from Qwen3-1.7B, a 1.7 billion parameter model with a 32k context length.
- Training Method: Utilizes GRPO, a technique introduced in the "DeepSeekMath" paper, which focuses on pushing the limits of mathematical reasoning in open language models.
- Framework: Trained using the TRL library (Transformers Reinforcement Learning).
Potential Use Cases
- Mathematical Reasoning: Ideal for applications requiring enhanced logical and mathematical problem-solving abilities.
- Code Generation: While not explicitly stated, models fine-tuned with reasoning-focused methods can often show improvements in structured output tasks like code generation.
- Research: Useful for researchers exploring the impact of GRPO on smaller language models and its effectiveness in improving specific reasoning skills.