ssurface/qwen3-4b-gdpo-length-sft-l3

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 1, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The ssurface/qwen3-4b-gdpo-length-sft-l3 model is a 4 billion parameter Qwen3-based language model, fine-tuned for compressed chain-of-thought reasoning at Level 3 (Symbolic). It leverages a training pipeline involving Supervised Fine-Tuning (SFT) followed by GRPO with a new reward mechanism. This model is specifically designed to excel in symbolic reasoning tasks, providing structured and efficient problem-solving capabilities within its 32K context length.

Loading preview...

Qwen3-4B GRPO - Level 3 (Symbolic) Overview

This model, ssurface/qwen3-4b-gdpo-length-sft-l3, is a 4 billion parameter variant of the Qwen3-4B-Instruct architecture. It has undergone a specialized fine-tuning process to enhance its capabilities in compressed chain-of-thought reasoning, specifically targeting Level 3 (Symbolic) tasks. This makes it particularly adept at handling problems that require symbolic manipulation and structured logical deduction.

Key Training Details

The model's development involved a multi-stage pipeline:

  • Initial Base: Started from Qwen/Qwen3-4B-Instruct-2507.
  • Supervised Fine-Tuning (SFT): Applied LoRA fine-tuning using ssurface/qwen3-4b-cot-compress-l3.
  • GRPO Optimization: Further fine-tuned with GRPO (Generalized Reinforcement Learning from Human Feedback) incorporating a novel reward function to optimize for reasoning efficiency and quality.

Primary Use Case

This model is designed for applications requiring advanced symbolic reasoning. Users can prompt it with problems explicitly requesting "Level 3 (Symbolic)" solutions, making it suitable for tasks where structured, step-by-step logical inference is crucial. Its training focuses on generating concise yet comprehensive reasoning paths.