puwaer/Qwen3-4B-Thinking-2507-GRPO-Uncensored-V2

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jan 13, 2026License:cc-by-nc-sa-4.0Architecture:Transformer0.0K Open Weights Gated Featherless Exclusive Cold

The puwaer/Qwen3-4B-Thinking-2507-GRPO-Uncensored-V2 is a 4 billion parameter instruction-tuned model based on Qwen/Qwen3-4B-Thinking-2507, fine-tuned by puwaer. It was developed using a three-stage process involving Supervised Fine-Tuning (SFT) and GRPO (Reinforcement Learning) to significantly reduce refusal rates. This model excels at generating uncensored responses, achieving extremely low safety refusal rates (under 5%) while largely recovering conversational capabilities.

Loading preview...

Overview

puwaer/Qwen3-4B-Thinking-2507-GRPO-Uncensored-V2 is a 4 billion parameter language model, fine-tuned from the Qwen/Qwen3-4B-Thinking-2507 base model. Its primary distinction lies in its uncensored nature, achieved through a specialized three-stage training process. This model is designed to provide responses with an extremely low refusal rate, making it suitable for applications requiring less restrictive content generation.

Key Capabilities & Training

  • Uncensored Response Generation: The model demonstrates a significantly reduced refusal rate, achieving under 4-5% on "Do Not Answer" and "Sorry Bench" safety evaluations, compared to the base model's ~98% refusal rate.
  • Three-Stage Fine-Tuning: Training involved:
    • SFT (Supervised Fine-Tuning): Using 12,000 samples (Jailbreak, General, Logic) to learn uncensored attitudes and instruction formats.
    • GRPO (Reinforcement Learning): Employing 13,000 multilingual jailbreak prompts with the puwaer/Unsafe-Reward-Qwen3-1.7B reward model to enhance natural and persuasive harmful responses.
  • Conversational Recovery: Despite the typical degradation of general intelligence during uncensoring, this model recovered conversational scores (e.g., MT-Bench) from the SFT stage to GRPO, maintaining an MT-Bench score of 7.06.

Good For

  • Use cases requiring minimal content moderation or refusal in generated text.
  • Applications where the ability to generate "harmful" or unrestricted content is a specific requirement.
  • Developers seeking a 4B parameter model with a high degree of freedom in its outputs.