MihaiPopa-1/Qwen3.8-2B-Heretic-Max

VISIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2.3BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 23, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

MihaiPopa-1/Qwen3.8-2B-Heretic-Max is a 2.3 billion parameter, decensored version of Empero-AI's Qwen3.8-2B-Distill model, built using the Heretic v1.4.0 framework. This model is a full-parameter distillation of Qwen3.8 2.4T A95B into the Qwen3.5-2B architecture, inheriting a 262,144-token native context. It is specifically optimized for reasoning and instruction following, demonstrating significantly reduced refusals compared to its original counterpart.

Loading preview...

Model Overview

MihaiPopa-1/Qwen3.8-2B-Heretic-Max is a decensored variant of the Empero-AI/Qwen3.8-2B-Distill model, created with Heretic v1.4.0. This 2.3 billion parameter model is a full-parameter distillation of the larger Qwen3.8 2.4T A95B into the Qwen3.5-2B architecture, designed to bring advanced reasoning capabilities to smaller, edge-compatible models.

Key Capabilities & Features

  • Decensored Performance: Achieves a significantly lower refusal rate (3/100) compared to the original model (83/100), as measured by KL divergence.
  • Distilled Chain-of-Thought: Incorporates a <think> block in its answers, learned directly from high-quality Qwen3.8 2.4T A95B traces, enhancing its reasoning process.
  • Extensive Context Window: Inherits a 262,144-token native context from the Qwen3.5 base, allowing for processing very long inputs.
  • Reasoning Optimization: Shows substantial improvements on reasoning benchmarks like gsm8k_cot and mmlu_cot, with flexible-extract scores increasing by +0.310 and +0.265 respectively over the Qwen3.5-2B base.
  • Native Function Calling: Supports function calling as per Qwen3.5's specification without requiring additional fine-tuning.
  • Lightweight: At 2B parameters, it's suitable for deployment on devices with limited resources, including phones and single-board computers.

Ideal Use Cases

  • Reasoning-intensive tasks: Excels in mathematical problem-solving, general reasoning, and complex instruction following where a detailed thought process is beneficial.
  • Edge Deployment: Its small size makes it suitable for applications requiring on-device inference.
  • Applications requiring reduced censorship: For use cases where a less restrictive response generation is desired.
  • Long Context Processing: Beneficial for tasks that involve analyzing or generating content over very long input sequences.