clijo/qwen3-4b-instruct-2507-bf16-reco-grpo-b200-clear-indigo-willow

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 20, 2026Architecture:Transformer Featherless Exclusive Cold

The clijo/qwen3-4b-instruct-2507-bf16-reco-grpo-b200-clear-indigo-willow model is a 4 billion parameter instruction-tuned causal language model, fine-tuned from Qwen/Qwen3-4B-Instruct-2507. It was trained using the GRPO method, which is designed to enhance mathematical reasoning capabilities. This model is particularly suited for tasks requiring improved logical and mathematical problem-solving, building upon the base Qwen3 architecture.

Loading preview...

Overview

This model, clijo/qwen3-4b-instruct-2507-bf16-reco-grpo-b200-clear-indigo-willow, is a 4 billion parameter instruction-tuned variant of the Qwen3-4B-Instruct-2507 base model. It has been specifically fine-tuned using the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). The training was conducted using the TRL framework.

Key Capabilities

  • Enhanced Mathematical Reasoning: The application of the GRPO method suggests an optimization for tasks requiring mathematical and logical problem-solving.
  • Instruction Following: As an instruction-tuned model, it is designed to follow user prompts and generate relevant responses.
  • Qwen3 Architecture: Benefits from the foundational capabilities of the Qwen3 model family.

Use Cases

  • Mathematical Problem Solving: Ideal for applications that involve numerical reasoning, equations, or logical deductions.
  • General Instruction Following: Suitable for a wide range of conversational AI and text generation tasks where precise instruction adherence is important.
  • Research and Development: Can serve as a base for further fine-tuning on specialized mathematical or reasoning datasets.