ssurface/qwen3-4b-gdpo-length-sft-l3
The ssurface/qwen3-4b-gdpo-length-sft-l3 model is a 4 billion parameter Qwen3-based language model, fine-tuned for compressed chain-of-thought reasoning at Level 3 (Symbolic). It leverages a training pipeline involving Supervised Fine-Tuning (SFT) followed by GRPO with a new reward mechanism. This model is specifically designed to excel in symbolic reasoning tasks, providing structured and efficient problem-solving capabilities within its 32K context length.
Loading preview...
Qwen3-4B GRPO - Level 3 (Symbolic) Overview
This model, ssurface/qwen3-4b-gdpo-length-sft-l3, is a 4 billion parameter variant of the Qwen3-4B-Instruct architecture. It has undergone a specialized fine-tuning process to enhance its capabilities in compressed chain-of-thought reasoning, specifically targeting Level 3 (Symbolic) tasks. This makes it particularly adept at handling problems that require symbolic manipulation and structured logical deduction.
Key Training Details
The model's development involved a multi-stage pipeline:
- Initial Base: Started from
Qwen/Qwen3-4B-Instruct-2507. - Supervised Fine-Tuning (SFT): Applied LoRA fine-tuning using
ssurface/qwen3-4b-cot-compress-l3. - GRPO Optimization: Further fine-tuned with GRPO (Generalized Reinforcement Learning from Human Feedback) incorporating a novel reward function to optimize for reasoning efficiency and quality.
Primary Use Case
This model is designed for applications requiring advanced symbolic reasoning. Users can prompt it with problems explicitly requesting "Level 3 (Symbolic)" solutions, making it suitable for tasks where structured, step-by-step logical inference is crucial. Its training focuses on generating concise yet comprehensive reasoning paths.