MihaiPopa-1/Qwen3.8-2B-Heretic-Max
MihaiPopa-1/Qwen3.8-2B-Heretic-Max is a 2.3 billion parameter, decensored version of Empero-AI's Qwen3.8-2B-Distill model, built using the Heretic v1.4.0 framework. This model is a full-parameter distillation of Qwen3.8 2.4T A95B into the Qwen3.5-2B architecture, inheriting a 262,144-token native context. It is specifically optimized for reasoning and instruction following, demonstrating significantly reduced refusals compared to its original counterpart.
Loading preview...
Model Overview
MihaiPopa-1/Qwen3.8-2B-Heretic-Max is a decensored variant of the Empero-AI/Qwen3.8-2B-Distill model, created with Heretic v1.4.0. This 2.3 billion parameter model is a full-parameter distillation of the larger Qwen3.8 2.4T A95B into the Qwen3.5-2B architecture, designed to bring advanced reasoning capabilities to smaller, edge-compatible models.
Key Capabilities & Features
- Decensored Performance: Achieves a significantly lower refusal rate (3/100) compared to the original model (83/100), as measured by KL divergence.
- Distilled Chain-of-Thought: Incorporates a
<think>block in its answers, learned directly from high-quality Qwen3.8 2.4T A95B traces, enhancing its reasoning process. - Extensive Context Window: Inherits a 262,144-token native context from the Qwen3.5 base, allowing for processing very long inputs.
- Reasoning Optimization: Shows substantial improvements on reasoning benchmarks like
gsm8k_cotandmmlu_cot, with flexible-extract scores increasing by +0.310 and +0.265 respectively over the Qwen3.5-2B base. - Native Function Calling: Supports function calling as per Qwen3.5's specification without requiring additional fine-tuning.
- Lightweight: At 2B parameters, it's suitable for deployment on devices with limited resources, including phones and single-board computers.
Ideal Use Cases
- Reasoning-intensive tasks: Excels in mathematical problem-solving, general reasoning, and complex instruction following where a detailed thought process is beneficial.
- Edge Deployment: Its small size makes it suitable for applications requiring on-device inference.
- Applications requiring reduced censorship: For use cases where a less restrictive response generation is desired.
- Long Context Processing: Beneficial for tasks that involve analyzing or generating content over very long input sequences.