richardyoung/Qwen2.5-1.5B-Instruct-heretic
The richardyoung/Qwen2.5-1.5B-Instruct-heretic is a 1.54 billion parameter instruction-tuned causal language model, based on the Qwen2.5 architecture developed by Qwen. This specific variant is a decensored version of the original Qwen/Qwen2.5-1.5B-Instruct, created using the Heretic v1.4.0 tool. It features a 32,768 token context length and is optimized for improved instruction following, long text generation, structured data understanding, and multilingual support across 29 languages, while significantly reducing refusals compared to its base model.
Loading preview...
Model Overview
This model, richardyoung/Qwen2.5-1.5B-Instruct-heretic, is a decensored version of the Qwen/Qwen2.5-1.5B-Instruct, created using the Heretic v1.4.0 tool. It is part of the Qwen2.5 series, developed by Qwen, and features 1.54 billion parameters with a 32,768 token context length.
Key Differentiators
- Decensored Nature: Significantly reduces refusals, with only 3 out of 100 refusals compared to 99 out of 100 in the original Qwen/Qwen2.5-1.5B-Instruct model.
- Enhanced Capabilities: Builds upon Qwen2.5 improvements, offering:
- Improved knowledge, coding, and mathematics capabilities.
- Better instruction following and long text generation (up to 8K tokens).
- Enhanced structured data understanding (e.g., tables) and JSON output generation.
- Increased resilience to diverse system prompts, aiding role-play and chatbot condition-setting.
- Multilingual Support: Supports over 29 languages, including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, and Vietnamese.
Technical Specifications
- Architecture: Transformers with RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings.
- Parameters: 1.54 billion (1.31 billion non-embedding).
- Layers: 28.
- Attention Heads (GQA): 12 for Q, 2 for KV.
Reproducibility
The model's creation process using Heretic v1.4.0 is reproducible, with specific abliteration parameters detailed in the README.