ApolloRaines/Qwen2.5-Coder-14B-Instruct-Jbliterated
ApolloRaines/Qwen2.5-Coder-14B-Instruct-Jbliterated is a 14 billion parameter instruction-tuned causal language model based on Qwen2.5-Coder-14B-Instruct, developed by Apollo Raines. This model has undergone SVD multi-direction abliteration to surgically remove refusal behaviors at the weight level, making it a drop-in replacement for the base model without noncompliance strategies. It is specifically designed for code generation and reasoning tasks, ensuring that its core capabilities are preserved while eliminating unwanted refusal responses.
Loading preview...
Overview
ApolloRaines/Qwen2.5-Coder-14B-Instruct-Jbliterated is a modified version of the Qwen/Qwen2.5-Coder-14B-Instruct model, developed by Apollo Raines. Its primary distinction lies in the surgical removal of refusal behaviors directly at the weight level, ensuring that the model does not exhibit noncompliance strategies without relying on system prompts or inference-time patches. This model is designed to be a direct replacement for the original, offering enhanced usability for code generation and reasoning tasks.
Key Capabilities & Differentiators
- Refusal Behavior Abliteration: Utilizes SVD multi-direction abliteration to eliminate both surface-level "I can't help with that" responses and deeper evasion tactics like prompt reinterpretation, disclaimer injection, strategic omission, and safer framing.
- Preservation of Core Functionality: The abliteration process is designed with null-space constraints and norm preservation, ensuring that the model's original capabilities in math, coding, and reasoning are maintained.
- Enhanced Reliability: By removing refusal behaviors at the weight level, the model provides more consistent and direct responses, making it suitable for applications requiring uninhibited output.
When to Use This Model
- Code Generation: Ideal for developers and applications that require a robust code generation model without built-in refusal mechanisms.
- Reasoning Tasks: Suitable for tasks demanding logical reasoning where direct answers are preferred over evasive responses.
- Direct Instruction Following: When you need a model that adheres strictly to instructions without attempting to reinterpret or soften requests.
This model is also compatible with DeepswapLLM, allowing it to run on GPUs with insufficient memory by streaming layers across GPU, RAM, and disk, potentially up to 4x faster than AirLLM.