Ksgk-fy/forgetting-gap-qwen3-4b-s0

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 12, 2026Architecture:Transformer Featherless Exclusive Cold

Ksgk-fy/forgetting-gap-qwen3-4b-s0 is a 4 billion parameter language model fine-tuned from Qwen/Qwen3-4B, utilizing the TRL framework. This model was specifically trained with the GRPO method, as introduced in the DeepSeekMath paper, to enhance its mathematical reasoning capabilities. With a context length of 32768 tokens, it is designed for tasks requiring advanced logical and mathematical problem-solving.

Loading preview...

Overview

Ksgk-fy/forgetting-gap-qwen3-4b-s0 is a 4 billion parameter language model derived from the Qwen/Qwen3-4B base model. It has been fine-tuned using the TRL (Transformers Reinforcement Learning) framework.

Key Differentiator: GRPO Training

This model's primary distinction lies in its training methodology. It was trained with GRPO (Gradient-based Reward Policy Optimization), a method detailed in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). This specialized training aims to significantly improve the model's proficiency in mathematical reasoning tasks.

Capabilities

  • Enhanced Mathematical Reasoning: Optimized through GRPO for complex mathematical problem-solving.
  • Large Context Window: Supports a context length of 32768 tokens, allowing for processing extensive inputs.
  • Qwen3-4B Foundation: Benefits from the robust architecture and general language understanding of the Qwen3-4B base model.

Use Cases

This model is particularly well-suited for applications requiring strong logical and mathematical inference, such as:

  • Solving mathematical word problems.
  • Assisting in scientific computations and derivations.
  • Tasks where precise reasoning and numerical accuracy are critical.