sergiopaniego/qwen3-0.6b-mbpp-grpo-k16

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 24, 2026Architecture:Transformer Featherless Exclusive Cold

The sergiopaniego/qwen3-0.6b-mbpp-grpo-k16 model is a 0.8 billion parameter language model, fine-tuned from Qwen/Qwen3-0.6B. It has been specifically trained on the google-research-datasets/mbpp dataset using the GRPO method, which is designed to enhance mathematical reasoning. This model is optimized for code generation and problem-solving tasks, particularly those involving programming challenges. It offers a 32768 token context length, making it suitable for handling moderately long code snippets and related instructions.

Loading preview...

Model Overview

This model, sergiopaniego/qwen3-0.6b-mbpp-grpo-k16, is a specialized version of the Qwen3-0.6B base model, developed by sergiopaniego. It features 0.8 billion parameters and supports a 32768 token context length.

Key Capabilities

  • Code Generation: Fine-tuned specifically on the google-research-datasets/mbpp dataset, which focuses on programming problems.
  • Enhanced Reasoning: Utilizes the GRPO (Gradient-based Reward Policy Optimization) training method, as introduced in the DeepSeekMath paper, to improve mathematical and logical reasoning capabilities, particularly relevant for code-related tasks.
  • TRL Framework: Trained using the TRL library, indicating a reinforcement learning approach to fine-tuning.

Ideal Use Cases

This model is particularly well-suited for applications requiring:

  • Solving programming challenges and generating code snippets.
  • Assisting with mathematical reasoning within a coding context.
  • Tasks that benefit from a model fine-tuned on a dataset of programming problems.