sergiopaniego/qwen3-0.6b-mbpp-grpo-k8

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 24, 2026Architecture:Transformer Featherless Exclusive Cold

The sergiopaniego/qwen3-0.6b-mbpp-grpo-k8 model is a 0.8 billion parameter Qwen3-0.6B variant fine-tuned by sergiopaniego. It was trained on the MBPP dataset using the GRPO method, which is designed to enhance mathematical reasoning. This model is optimized for code generation and problem-solving tasks, particularly those involving mathematical or logical structures.

Loading preview...

Model Overview

This model, sergiopaniego/qwen3-0.6b-mbpp-grpo-k8, is a fine-tuned version of the Qwen3-0.6B architecture, developed by sergiopaniego. It has approximately 0.8 billion parameters and a context length of 32768 tokens.

Key Capabilities

  • Code Generation and Problem Solving: The model was specifically fine-tuned on the MBPP (Mostly Basic Python Problems) dataset, indicating a strong focus on generating and solving programming-related tasks.
  • Enhanced Mathematical Reasoning: Training utilized the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the DeepSeekMath paper. This method is designed to improve a model's capabilities in mathematical reasoning.

Training Details

The model was trained using the TRL (Transformers Reinforcement Learning) library. The application of GRPO suggests an emphasis on optimizing performance for tasks requiring logical and mathematical understanding, making it distinct from general-purpose instruction-tuned models.

Good For

  • Developers and researchers working on code generation.
  • Applications requiring mathematical problem-solving or logical reasoning within a coding context.
  • Experimentation with GRPO-trained models for specific task improvements.