ssurface/qwen3-4b-gdpo-length-sft-l2
The ssurface/qwen3-4b-gdpo-length-sft-l2 model is a 4 billion parameter Qwen3-based causal language model developed by ssurface, fine-tuned for compressed chain-of-thought reasoning at a 'Level 2 (Concise)' output. It utilizes a SFT then GRPO training pipeline with a new reward mechanism to optimize for concise reasoning. This model is specifically designed for tasks requiring efficient and condensed problem-solving explanations within its 32768 token context length.
Loading preview...
Model Overview
The ssurface/qwen3-4b-gdpo-length-sft-l2 is a 4 billion parameter model built upon the Qwen3-4B-Instruct architecture. It has undergone a specialized fine-tuning process to excel in compressed chain-of-thought reasoning, specifically targeting a 'Level 2 (Concise)' output style.
Key Capabilities
- Concise Chain-of-Thought Reasoning: Optimized to provide condensed and efficient reasoning steps, rather than verbose explanations.
- GRPO Fine-tuning: Leverages a SFT (Supervised Fine-Tuning) followed by GRPO (Generalized Reinforcement Learning from Human Feedback) with a novel reward function to achieve its concise output style.
- Qwen3 Base: Benefits from the foundational capabilities of the Qwen3-4B-Instruct model.
Training Pipeline
The model's unique capabilities stem from its multi-stage training:
- Base Model: Started with
Qwen/Qwen3-4B-Instruct-2507. - SFT LoRA: Fine-tuned using LoRA (Low-Rank Adaptation) with
ssurface/qwen3-4b-cot-compress-l2. - GRPO with New Reward: Further optimized using GRPO with a custom reward mechanism to reinforce concise reasoning.
Good For
- Applications requiring efficient and brief reasoning outputs.
- Scenarios where detailed, step-by-step explanations need to be condensed.
- Tasks that benefit from a 'Level 2 (Concise)' problem-solving approach.