richardyoung/Qwen2.5-3B-Instruct-heretic
The richardyoung/Qwen2.5-3B-Instruct-heretic is a 3.09 billion parameter instruction-tuned causal language model, based on the Qwen2.5 architecture developed by Qwen. This model is a decensored version of the original Qwen2.5-3B-Instruct, created using the Heretic v1.4.0 tool, significantly reducing refusals from 96/100 to 3/100. It features improved capabilities in coding, mathematics, instruction following, and generating structured outputs like JSON, with a full context length of 32,768 tokens.
Loading preview...
Overview
This model, richardyoung/Qwen2.5-3B-Instruct-heretic, is a 3.09 billion parameter instruction-tuned causal language model derived from the Qwen2.5 series by Qwen. It has been processed with Heretic v1.4.0 to create a "decensored" version of the original Qwen/Qwen2.5-3B-Instruct.
Key Differentiators
- Decensored Output: Significantly reduces refusal rates from 96/100 in the original model to 3/100, making it more permissive in its responses.
- Enhanced Core Capabilities: Builds upon Qwen2.5's improvements in coding, mathematics, and instruction following.
- Structured Data Handling: Excels at understanding structured data (e.g., tables) and generating structured outputs, particularly JSON.
- Long Context Support: Supports a full context length of 32,768 tokens and can generate up to 8,192 tokens.
- Multilingual: Offers support for over 29 languages.
Performance
Compared to the original Qwen2.5-3B-Instruct, this 'heretic' version shows a KL divergence of 0.0494, indicating a measurable shift in its output distribution, primarily reflected in its reduced refusal rate.
Good for
- Applications requiring a less restrictive instruction-tuned model.
- Tasks involving code generation and mathematical problem-solving.
- Generating structured outputs, such as JSON, and processing tabular data.
- Chatbot implementations that require robust role-play and condition-setting capabilities.