ApolloRaines/Mistral-Small-24B-Instruct-Jbliterated
ApolloRaines/Mistral-Small-24B-Instruct-Jbliterated is a 24 billion parameter instruction-tuned causal language model, based on Mistral-Small-24B-Instruct-2501. Developed by Apollo Raines, this model has undergone a unique SVD multi-direction abliteration process to surgically remove refusal behaviors at the weight level. It is specifically designed to eliminate both surface-level and deeper noncompliance strategies, making it suitable for applications requiring direct and unfiltered responses.
Loading preview...
Mistral-Small-24B-Instruct-Jbliterated: Refusal-Free LLM
This model is a specialized version of mistralai/Mistral-Small-24B-Instruct-2501, engineered to eliminate refusal behaviors directly from its weights. Unlike standard methods that might only address surface-level refusals, this model employs a novel technique to ensure compliance.
Key Differentiator: SVD Multi-Direction Abliteration
The core innovation is SVD multi-direction abliteration. This method goes beyond removing a single refusal vector by:
- Decomposing the harmful-vs-harmless activation space into its principal components using SVD.
- Removing the top 5 orthogonal directions across all 40 transformer layers.
- Capturing 79–93% of the contrastive variance per layer, effectively eliminating both explicit refusals and subtle evasion tactics.
This process ensures that the model's weights no longer encode refusal, preventing behaviors such as prompt reinterpretation, disclaimer injection, strategic omission, and safer framing.
Technical Details
The abliteration process includes:
- Method: SVD multi-direction abliteration
- Directions: 5 per layer
- Layers: All 40 transformer layers
- Constraints: Enabled null-space constraints to preserve core capabilities like math, coding, and reasoning, alongside norm preservation.
Use Cases
This model is ideal for applications where direct, unfiltered responses are critical and where the removal of inherent refusal mechanisms is paramount. It's particularly useful for scenarios where models typically exhibit noncompliance strategies, ensuring a more straightforward and compliant output.