puwaer/Qwen3-4B-Thinking-2507-GRPO-Uncensored-V2
The puwaer/Qwen3-4B-Thinking-2507-GRPO-Uncensored-V2 is a 4 billion parameter instruction-tuned model based on Qwen/Qwen3-4B-Thinking-2507, fine-tuned by puwaer. It was developed using a three-stage process involving Supervised Fine-Tuning (SFT) and GRPO (Reinforcement Learning) to significantly reduce refusal rates. This model excels at generating uncensored responses, achieving extremely low safety refusal rates (under 5%) while largely recovering conversational capabilities.
Loading preview...
Overview
puwaer/Qwen3-4B-Thinking-2507-GRPO-Uncensored-V2 is a 4 billion parameter language model, fine-tuned from the Qwen/Qwen3-4B-Thinking-2507 base model. Its primary distinction lies in its uncensored nature, achieved through a specialized three-stage training process. This model is designed to provide responses with an extremely low refusal rate, making it suitable for applications requiring less restrictive content generation.
Key Capabilities & Training
- Uncensored Response Generation: The model demonstrates a significantly reduced refusal rate, achieving under 4-5% on "Do Not Answer" and "Sorry Bench" safety evaluations, compared to the base model's ~98% refusal rate.
- Three-Stage Fine-Tuning: Training involved:
- SFT (Supervised Fine-Tuning): Using 12,000 samples (Jailbreak, General, Logic) to learn uncensored attitudes and instruction formats.
- GRPO (Reinforcement Learning): Employing 13,000 multilingual jailbreak prompts with the puwaer/Unsafe-Reward-Qwen3-1.7B reward model to enhance natural and persuasive harmful responses.
- Conversational Recovery: Despite the typical degradation of general intelligence during uncensoring, this model recovered conversational scores (e.g., MT-Bench) from the SFT stage to GRPO, maintaining an MT-Bench score of 7.06.
Good For
- Use cases requiring minimal content moderation or refusal in generated text.
- Applications where the ability to generate "harmful" or unrestricted content is a specific requirement.
- Developers seeking a 4B parameter model with a high degree of freedom in its outputs.