htb-ac-1424625/Qwen3.8-2B-Distill-heretic
The htb-ac-1424625/Qwen3.8-2B-Distill-heretic model is a 2.3 billion parameter causal language model developed by Empero, based on the Qwen3.5-2B architecture. It is a decensored version of the Qwen3.8-2B-Distill model, created using Heretic, and features a 262,144-token native context length. This model is specifically distilled for reasoning and instruction following, learning directly from Qwen3.8 2.4T A95B traces, making it highly optimized for mathematical and general reasoning tasks, even on edge devices.
Loading preview...
Overview
This model, htb-ac-1424625/Qwen3.8-2B-Distill-heretic, is a 2.3 billion parameter causal language model developed by Empero. It is a decensored variant of the empero-ai/Qwen3.8-2B-Distill model, created using the Heretic tool. Built upon the Qwen3.5-2B architecture, it features a substantial 262,144-token native context length. The model is a full-parameter distillation, meaning every parameter was updated during training, rather than using an adapter.
Key Capabilities
- Distilled Chain-of-Thought: Learns directly from high-quality Qwen3.8 2.4T A95B traces, enabling it to generate reasoning steps (e.g.,
<think>blocks) for complex problems. - Reasoning and Instruction Following: Trained on the same curriculum as its larger 4B and 9B siblings, focusing on mathematics, general reasoning, and instruction following.
- Edge Device Compatibility: Its 2B parameter size allows it to run efficiently on devices with limited resources, such as phones, single-board computers, and CPU-only machines, especially with quantized builds.
- Native Function Calling: Supports function calling as per Qwen3.5's specification without requiring additional wrappers or fine-tuning.
- Reduced Refusals: Demonstrates significantly fewer refusals (3/100) compared to its original counterpart (84/100).
Good For
- Mathematical and General Reasoning Tasks: Excels in tasks requiring logical deduction and problem-solving, as evidenced by significant performance gains on
gsm8k_cotandmmlu_cotbenchmarks compared to its base model. - Instruction Following: Designed to accurately follow complex instructions.
- Resource-Constrained Environments: Ideal for deployment on edge devices due to its small size and efficient architecture.
- Applications Requiring Decensored Output: Suitable for use cases where unfiltered responses are preferred or necessary.
- Developers Seeking Reproducibility: The model's creation process using Heretic is reproducible, with details provided in the
reproducedirectory.