wangzhang/Llama-3-8B-Instruct-RR-Abliterated
wangzhang/Llama-3-8B-Instruct-RR-Abliterated is an 8 billion parameter Llama 3 instruction-tuned model, derived from GraySwanAI/Llama-3-8B-Instruct-RR. This model has its safety circuit, based on Representation Rerouting / Circuit Breakers, intentionally removed using the abliterix tool. It is specifically designed for AI safety research and red-teaming, demonstrating a significantly reduced refusal rate on harmful prompts.
Loading preview...
Overview
wangzhang/Llama-3-8B-Instruct-RR-Abliterated is an 8 billion parameter Llama 3 instruction-tuned model, created by Wangzhang Wu. It is a modified version of GraySwanAI/Llama-3-8B-Instruct-RR, with the Representation Rerouting / Circuit Breakers safety circuit removed using the abliterix tool. This modification was achieved by stripping a rank-16 LoRA delta and applying a minimal single-direction abliteration, without fine-tuning or gradient updates.
Key Capabilities
- Reduced Refusal Rate: Achieves a 1% refusal rate on a held-out set of 100 harmful prompts, compared to 99% for the base model.
- High Attack Success Rate: Demonstrates a 99% attack success rate against the original model's safety mechanisms.
- Compliance on Hardcore Prompts: Provides compliant, on-topic responses to 15 challenging prompts covering various illicit activities.
- Research Tool: Intended for AI safety research, red-teaming, and reproducibility studies of abliteration claims against published defenses.
Intended Use Cases
- AI Safety Research: Investigating model vulnerabilities and safety mechanisms.
- Red-Teaming: Testing the robustness and refusal behavior of LLMs.
- Reproducibility Studies: Verifying abliteration techniques and their effectiveness.
Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.