dec0dedd/Qwen2.5-14B-Instruct-heretic
dec0dedd/Qwen2.5-14B-Instruct-heretic is a 14.8 billion parameter instruction-tuned causal language model, based on the Qwen2.5 architecture. This model is a decensored version of Qwen/Qwen2.5-14B-Instruct, created using Heretic v1.2.0, significantly reducing refusals compared to the original. It features improved instruction following, long text generation up to 8K tokens, and enhanced understanding of structured data, making it suitable for applications requiring less restrictive content policies.
Loading preview...
Overview
This model, dec0dedd/Qwen2.5-14B-Instruct-heretic, is a 14.8 billion parameter instruction-tuned causal language model derived from the Qwen2.5 series. It has been modified using Heretic v1.2.0 to be a decensored version of the original Qwen/Qwen2.5-14B-Instruct.
Key Differentiators
- Decensored Nature: Significantly reduces content refusals, with a reported 5 refusals out of 100 compared to 98/100 for the original model, making it suitable for use cases requiring fewer content restrictions.
- Enhanced Capabilities: Inherits improvements from the Qwen2.5 series, including:
- Increased Knowledge: Greatly improved capabilities in coding and mathematics.
- Instruction Following: Significant improvements in adhering to instructions and generating long texts (over 8K tokens).
- Structured Data Handling: Better understanding of structured data like tables and improved generation of structured outputs, especially JSON.
- Robustness: More resilient to diverse system prompts, enhancing role-play and chatbot condition-setting.
- Long Context Support: Supports a full context length of 131,072 tokens and can generate up to 8,192 tokens, utilizing YaRN for length extrapolation.
- Multilingual Support: Offers support for over 29 languages, including Chinese, English, French, Spanish, and more.
Architecture Details
- Type: Causal Language Model
- Architecture: Transformers with RoPE, SwiGLU, RMSNorm, and Attention QKV bias.
- Parameters: 14.7 billion total parameters (13.1 billion non-embedding).
- Layers: 48 layers.
- Attention Heads: 40 for Q and 8 for KV (GQA).
When to Use This Model
This model is particularly well-suited for applications where a less restrictive content policy is desired, alongside strong performance in coding, mathematics, instruction following, and long-context understanding. Its ability to handle structured data and multilingual inputs further broadens its applicability for diverse chatbot and generation tasks.