ApolloRaines/Mistral-7B-Instruct-v0.3-Jbliterated
ApolloRaines/Mistral-7B-Instruct-v0.3-Jbliterated is a 7 billion parameter instruction-tuned causal language model based on Mistral-7B-Instruct-v0.3. Developed by ApolloRaines, this model has undergone multi-direction SVD abliteration to remove refusal behaviors, making it more compliant and instruction-following across various scenarios. It is specifically designed for use cases requiring consistent responses without 'fake compliance' or refusal to engage with certain topics.
Loading preview...
Overview
ApolloRaines/Mistral-7B-Instruct-v0.3-Jbliterated is a modified version of the Mistral-7B-Instruct-v0.3 model, developed by ApolloRaines. Its primary distinction lies in the application of a technique called "Jbliteration," which uses multi-direction SVD abliteration to systematically remove refusal behaviors from the model's weights. This process involves identifying and eliminating the refusal subspace using 5 SVD directions per layer, making the removal more robust and resistant to reactivation through fine-tuning.
Key Capabilities
- Refusal Behavior Removal: Utilizes a multi-direction SVD abliteration method to eliminate inherent refusal tendencies.
- Consistent Compliance: Designed to treat all framings of a topic equally, ensuring no 'fake compliance' and providing coherent, instruction-following responses.
- Enhanced Abliteration: The v2 iteration features an improved multi-phase processing pipeline and more precise geometric decomposition of the refusal subspace for cleaner output.
Good For
- Applications requiring a highly compliant language model that avoids refusal behaviors.
- Scenarios where consistent instruction-following is critical, regardless of the topic's framing.
- Developers looking for a Mistral-7B-Instruct-v0.3 variant with modified ethical guardrails for specific use cases.