richardyoung/Qwen2.5-0.5B-Instruct-heretic

TEXT GENERATIONConcurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 24, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The richardyoung/Qwen2.5-0.5B-Instruct-heretic is a 0.49 billion parameter instruction-tuned causal language model, based on the Qwen2.5 architecture by Qwen, with a 32,768 token context length. This model is a decensored version of the original Qwen2.5-0.5B-Instruct, created using the Heretic v1.4.0 project. It significantly reduces refusals compared to the base model, making it suitable for applications requiring less restrictive content generation. The model retains the Qwen2.5 improvements in knowledge, coding, mathematics, instruction following, and structured output generation.

Loading preview...

richardyoung/Qwen2.5-0.5B-Instruct-heretic Overview

This model is a decensored version of the Qwen/Qwen2.5-0.5B-Instruct, developed using the Heretic v1.4.0 project. It is a 0.49 billion parameter instruction-tuned causal language model built on the Qwen2.5 architecture, featuring a 32,768 token context length and 8,192 token generation capacity. The primary differentiator of this 'heretic' version is its significantly reduced refusal rate, with only 3 refusals out of 100 compared to 91/100 for the original model, as measured by KL divergence of 0.0925.

Key Capabilities Inherited from Qwen2.5

  • Enhanced Knowledge & Reasoning: Improved capabilities in coding and mathematics due to specialized expert models.
  • Instruction Following: Significant improvements in adhering to instructions and generating long texts (over 8K tokens).
  • Structured Data & Output: Better understanding of structured data like tables and improved generation of structured outputs, especially JSON.
  • System Prompt Resilience: More robust to diverse system prompts, enhancing role-play and chatbot condition-setting.
  • Multilingual Support: Supports over 29 languages, including Chinese, English, French, Spanish, and more.

When to Use This Model

This model is particularly well-suited for use cases where a less restrictive content policy is desired, and the original Qwen2.5-0.5B-Instruct's refusal rate is too high. It maintains the core strengths of the Qwen2.5 series, making it a good choice for applications requiring strong instruction following, coding assistance, mathematical problem-solving, and structured output generation, especially in a multilingual context, while offering greater flexibility in response generation.