insraq/Qwen3.5-4B-EmperoAI-Qwen3.8-Distill-Heretic-Abliterated
insraq/Qwen3.5-4B-EmperoAI-Qwen3.8-Distill-Heretic-Abliterated is a 4.5 billion parameter causal language model, a decensored version of EmperoAI's Qwen3.8-4B. This model is a full-parameter distillation of a frontier-scale Qwen3.8 teacher into the Qwen3.5-4B architecture, specifically optimized for reasoning, mathematics, and instruction following. It features a 262,144-token native context and is designed to bring advanced reasoning capabilities to consumer hardware.
Loading preview...
Model Overview
This model, insraq/Qwen3.5-4B-EmperoAI-Qwen3.8-Distill-Heretic-Abliterated, is a 4.5 billion parameter causal language model based on the Qwen3.5-4B architecture. It is a decensored version of EmperoAI's Qwen3.8-4B, created using the Heretic v1.4.0 tool, with specific abliteration parameters applied to modify its behavior.
Key Capabilities & Features
- Distilled Reasoning: It is a full-parameter distillation of a larger Qwen3.8 teacher model, trained on approximately 45,000 curated teacher traces focusing on chain-of-thought reasoning, mathematics, and instruction following. This allows it to generate
<think>blocks for detailed reasoning. - Efficient Size: With 4.5 billion parameters, it is designed to run efficiently on consumer hardware, with bf16 weights fitting in around 8 GB of VRAM.
- Extended Context: Inherits a native 262,144-token context length from its Qwen3.5 base.
- Native Function Calling: Supports function calling as per Qwen3.5's specification without requiring additional wrappers or fine-tuning.
- Reduced Refusals: Compared to the original EmperoAI Qwen3.8-4B, this 'abliterated' version shows significantly fewer refusals (6/100 vs. 99/100).
Performance Highlights
Benchmarks show that while it has a slightly lower gsm8k_cot score than the base Qwen3.5-4B, it achieves a substantial improvement in mmlu (CoT, 57 subjects) scores, with a +0.199 increase in flexible-extract accuracy, indicating enhanced multi-task language understanding and reasoning capabilities.
Use Cases
This model is particularly well-suited for applications requiring strong reasoning, mathematical problem-solving, and complex instruction following, especially where a smaller, more efficient model is desired for deployment on consumer-grade hardware. Users should be aware that it is a decensored version.