shad0wcrawl3r/Qwen2.5-1.5B-heretic

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 8, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

shad0wcrawl3r/Qwen2.5-1.5B-heretic is a 1.54 billion parameter causal language model, based on the Qwen2.5 architecture developed by Qwen, with a 32,768 token context length. This model is a decensored version of the original Qwen2.5-1.5B, created using the Heretic v1.4.0 tool. It features improved capabilities in coding, mathematics, instruction following, and long text generation, making it suitable for applications requiring less restrictive content generation.

Loading preview...

Model Overview

This model, shad0wcrawl3r/Qwen2.5-1.5B-heretic, is a 1.54 billion parameter causal language model derived from the Qwen2.5 series by Qwen. It is specifically a decensored version of the original Qwen2.5-1.5B, created using the Heretic v1.4.0 tool. The base Qwen2.5 architecture features transformers with RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings, supporting a substantial context length of 32,768 tokens.

Key Capabilities & Improvements

Building upon the Qwen2.5 foundation, this model inherits significant enhancements over its predecessor, Qwen2, including:

  • Expanded Knowledge & Specialized Skills: Greatly improved capabilities in coding and mathematics, leveraging specialized expert models.
  • Enhanced Instruction Following: Better adherence to instructions and more resilient to diverse system prompts, aiding in role-play and chatbot condition-setting.
  • Long-Context & Structured Data Handling: Improved generation of long texts (up to 8K tokens) and understanding/generation of structured data like JSON.
  • Multilingual Support: Supports over 29 languages, including major global languages.

Decensoring & Performance

The "heretic" modification aims to reduce refusals compared to the original model. While the original Qwen2.5-1.5B had 2 refusals out of 100, this decensored version shows 1 refusal out of 100, indicating a less restrictive output behavior. The abliteration parameters used for this modification are detailed in the original README, ensuring reproducibility.

Usage Recommendation

As a base language model, it is not recommended for direct conversational use without further post-training such as Supervised Fine-Tuning (SFT), Reinforcement Learning from Human Feedback (RLHF), or continued pretraining. Developers can leverage its decensored nature for applications requiring more open-ended content generation.