llmfan46/Qwen3.5-27B-uncensored-heretic-v1
The llmfan46/Qwen3.5-27B-uncensored-heretic-v1 is a 27 billion parameter language model, a decensored version of Qwen/Qwen3.5-27B created using the Heretic v1.2.0 with Arbitrary-Rank Ablation (ARA) method. This model significantly reduces refusals by 92% (8/100 vs 95/100 for the original) while maintaining a low KL divergence of 0.0331, indicating preserved quality. It supports a 32768 token context length natively and is primarily designed for applications requiring less restrictive content generation without compromising core model capabilities.
Loading preview...
llmfan46/Qwen3.5-27B-uncensored-heretic-v1: Decensored Qwen3.5-27B
This model is a 27 billion parameter decensored version of the original Qwen/Qwen3.5-27B, developed by llmfan46 using the Heretic v1.2.0 framework with the Arbitrary-Rank Ablation (ARA) method. The primary goal of this modification is to drastically reduce content refusals while preserving the model's original quality.
Key Capabilities
- Significantly Reduced Refusals: Achieves a 92% reduction in refusals (8/100 compared to 95/100 in the original model), making it suitable for use cases requiring less restrictive content generation.
- Preserved Model Quality: Maintains a low KL divergence of 0.0331 from the original Qwen3.5-27B, indicating that the decensoring process has minimal impact on the model's core linguistic and reasoning abilities.
- Robust Performance: Benchmarks show comparable performance to the original Qwen3.5-27B across various tasks, including MMLU (85.97% accuracy), instruction following (IFEval 95.0%, IFBench 76.5%), and coding (SWE-bench Verified 72.4%, LiveCodeBench v6 80.7%).
- Multimodal Support: Inherits the multimodal capabilities of the base Qwen3.5 model, including unified vision-language foundation, efficient hybrid architecture, scalable RL generalization, and global linguistic coverage (201 languages).
- Extended Context Length: Natively supports a context length of 32,768 tokens, extensible up to 1,010,000 tokens using RoPE scaling techniques like YaRN.
Good for
- Applications requiring uncensored or less restricted content generation: Ideal for creative writing, role-playing, or research where the original model's refusal rates might be prohibitive.
- Maintaining high-quality outputs: Suitable for tasks where preserving the original model's performance across knowledge, reasoning, and coding benchmarks is crucial.
- Multimodal tasks: Can be used for vision-language tasks, including image and video understanding, general VQA, text recognition, and document understanding.
- Agentic workflows: Recommended for building agent applications, especially with Qwen-Agent and Qwen Code, leveraging its tool-calling capabilities.