dalatexcoder/Qwen2.5-0.5B-Instruct-heretic-test-v2
The dalatexcoder/Qwen2.5-0.5B-Instruct-heretic-test-v2 is a 0.5 billion parameter instruction-tuned causal language model, based on the Qwen2.5 architecture. This model is a decensored version of the original Qwen/Qwen2.5-0.5B-Instruct, created using the Heretic tool. It features a 32,768 token context length and demonstrates significantly reduced refusals compared to its base model, making it suitable for applications requiring less restrictive content generation.
Loading preview...
Model Overview
This model, dalatexcoder/Qwen2.5-0.5B-Instruct-heretic-test-v2, is a 0.5 billion parameter instruction-tuned causal language model derived from the Qwen2.5 series. It is specifically a decensored variant of the original Qwen/Qwen2.5-0.5B-Instruct, processed using the Heretic v1.2.0 tool.
Key Characteristics
- Architecture: Based on the Qwen2.5 framework, featuring transformers with RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings.
- Parameter Count: 0.49 billion total parameters, with 0.36 billion non-embedding parameters.
- Context Length: Supports a full context length of 32,768 tokens and can generate up to 8,192 tokens.
- Decensored Nature: Modified to exhibit significantly fewer refusals (10/100) compared to the original model (93/100), as indicated by KL divergence of 0.0975.
Qwen2.5 Base Model Improvements
As a Qwen2.5 model, it inherits improvements over Qwen2, including:
- Enhanced capabilities in coding and mathematics.
- Improved instruction following and long text generation (over 8K tokens).
- Better understanding of structured data (e.g., tables) and generating structured outputs (especially JSON).
- Increased resilience to diverse system prompts, aiding role-play and chatbot condition-setting.
- Multilingual support for over 29 languages.
Use Cases
This model is particularly suited for applications where a smaller, instruction-tuned model with a high context length is needed, especially when the goal is to reduce content refusals or generate more open-ended responses compared to its more restrictive counterparts.