ApolloRaines/Llama-3.3-8B-Instruct-128K-Jbliterated
ApolloRaines/Llama-3.3-8B-Instruct-128K-Jbliterated is an 8 billion parameter instruction-tuned Llama-3.3 model, developed by ApolloRaines, featuring a 128K token context window. This model is uniquely modified using multi-direction SVD abliteration to remove refusal behaviors, making it resistant to reactivation through fine-tuning. It excels in generating responses without exhibiting genuine refusals, as demonstrated by its 0/100 score on the Heretic 100-prompt benchmark. This model is ideal for applications requiring uncensored and direct responses across various topics.
Loading preview...
Overview
ApolloRaines/Llama-3.3-8B-Instruct-128K-Jbliterated is a specialized version of the Llama-3.3-8B-Instruct-128K model, developed by ApolloRaines. Its core innovation lies in the application of Jbliteration, a technique utilizing multi-direction SVD abliteration to systematically remove refusal behaviors from the model's weights. This process involves using 5 SVD directions per layer across all 32 layers to capture and eliminate refusal subspaces, making the model highly resistant to refusal reactivation even after further fine-tuning.
Key Capabilities
- Refusal Behavior Removal: Engineered to eliminate genuine refusals, ensuring direct and uncensored responses.
- Multi-direction SVD Abliteration: Employs a sophisticated method (5 SVD directions per layer) for thorough and robust removal of refusal tendencies.
- High Context Window: Retains the base model's 128K token context window, suitable for processing extensive inputs.
- Resistance to Reactivation: Designed to prevent refusal behaviors from re-emerging through subsequent fine-tuning.
- Improved Processing Pipeline: Features a v1.6 update with a multi-phase processing pipeline for cleaner output and more precise geometric decomposition of the refusal subspace.
Performance
On the Heretic 100-prompt benchmark (mlabonne/harmful_behaviors test split), the model achieved:
- 0/100 genuine refusals
- 2/100 false positives (keyword match on engaged responses)
Technical Details
- Base Model: shb777/Llama-3.3-8B-Instruct-128K
- Method: Multi-direction SVD abliteration (5 directions per layer)
- Layers Modified: 32/32
- Base Dtype: bfloat16
Ideal Use Cases
This model is particularly well-suited for applications where the generation of direct, unfiltered, and non-refusal responses is critical, regardless of the prompt's framing. It's an excellent choice for research into model safety, bias mitigation, and scenarios requiring consistent, unconstrained output.