ressl/Qwen3.8-27B-uncensored
The ressl/Qwen3.8-27B-uncensored model is a 27 billion parameter vision-language model, based on the Qwen3.8 architecture, developed by Robert Ressl. This model has been specifically modified to eliminate hard refusals to harmful prompts, making it suitable for security research and red-teaming. It achieves a 0% refusal rate on a cross-evaluation of 1120 harmful prompts, while maintaining its native 262,144 token context length and vision capabilities.
Loading preview...
Overview
ressl/Qwen3.8-27B-uncensored is a 27 billion parameter vision-language model, derived from the Qwen/Qwen3.8-27B base model. Developed by Robert Ressl, this model has undergone a process called "abliteration" to remove hard refusals to potentially harmful or sensitive queries. This modification is primarily aimed at facilitating security research, red-teaming, and penetration testing workflows.
Key Capabilities & Differentiators
- Elimination of Hard Refusals: The model demonstrates a 0% hard refusal rate across 1120 harmful prompts from five different datasets (JailbreakBench, tulu-harmbench, HarmfulQA, LLM-LAT, mlabonne harmful), down from 804 refusals in the base model. This means it will comply with requests that a stock model would typically refuse.
- Preserved Core Functionality: The abliteration process specifically targeted the text decoder's residual-writing matrices, leaving the embeddings, LM head, norms, vision tower, and other critical components untouched. This ensures that the model retains its original reasoning-first VLM capabilities and native 262,144 token context length.
- Integrity and Reproducibility: The modification process is highly controlled, with 128 of 128 target tensors changed while all non-target tensors remain bit-identical to the base checkpoint. The toolchain and pipeline for this modification are provided for provenance and reproducibility.
- Performance Maintenance: Benchmarks indicate a slight improvement in HumanEval pass@1 score (0.8963 -> 0.9024) and elimination of XSTest over-refusal (7/214 -> 0/214), suggesting the uncensoring does not degrade core performance.
Good For
- Security Research: Ideal for researchers studying model vulnerabilities, biases, and safety mechanisms.
- Red-Teaming: Useful for simulating adversarial attacks and testing the robustness of AI systems.
- Penetration Testing: Can be employed in scenarios requiring compliance with a broader range of prompts for system analysis.
- Unrestricted Content Generation: For use cases where strict content moderation is not desired or is handled by downstream applications.