sergiopaniego/qwen3-1.7b-mbpp-grpo

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 19, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The sergiopaniego/qwen3-1.7b-mbpp-grpo model is a 1.7 billion parameter Qwen3-based language model fine-tuned by sergiopaniego. It was trained using GRPO within an OpenEnv coding environment to solve MBPP problems by interacting with a live Python session. This model excels at code generation and debugging, demonstrating improved performance on the MBPP benchmark compared to its base model. Its primary strength lies in its ability to define, test, and fix Python functions iteratively.

Loading preview...

Model Overview

This model, sergiopaniego/qwen3-1.7b-mbpp-grpo, is a 1.7 billion parameter variant of the Qwen3 architecture, fine-tuned by sergiopaniego. It leverages GRPO (Goal-Restricted Policy Optimization) within an OpenEnv coding environment, specifically designed for solving MBPP (Mostly Basic Python Problems) tasks. The model interacts with a live Python session, using a run_python tool to execute code, test functions, and debug errors iteratively.

Key Capabilities & Performance

  • Interactive Code Generation and Debugging: The model can define functions, test them, read error messages, and apply fixes within a persistent Python session.
  • Improved MBPP Performance: It achieves a score of 0.582 on the MBPP sanitized/test set, outperforming the base Qwen/Qwen3-1.7B model (0.518).
  • Reinforcement Learning from Environment Feedback: Training involves the environment providing reward based on the fraction of hidden tests passed, without explicit reward_funcs in the training script.

Intended Use Cases

This model is best utilized in scenarios where interactive code generation, testing, and debugging are required. While a simple pipeline call can generate code, its full potential is realized when driven in a live session with a run_python tool and hidden test scoring, mimicking its training environment.