Hyukkyu/Llama-3.1-8B-RAQUEL-WMDP-Unlearn-IDK-GD-LoRA-v1
Hyukkyu/Llama-3.1-8B-RAQUEL-WMDP-Unlearn-IDK-GD-LoRA-v1 is an 8 billion parameter Llama-3.1 model specifically unlearned from the RAQUEL WMDP experiments. This model utilizes the IDK+GD method to replace forgotten answers with "I don't know" refusals, while retaining desired knowledge. It is optimized for selective knowledge unlearning and retention, making it suitable for applications requiring precise control over model responses to specific queries.
Loading preview...
Overview
This model, Hyukkyu/Llama-3.1-8B-RAQUEL-WMDP-Unlearn-IDK-GD-LoRA-v1, is an 8 billion parameter Llama-3.1 variant that has undergone a specialized unlearning process. It was derived from Hyukkyu/Llama-3.1-8B-RAQUEL-WMDP-M-orig-LoRA-v1 using the IDK+GD method, which aims to make the model refuse to answer specific "forget" questions by responding with "I don't know," while preserving its ability to answer "retain" questions.
Key Capabilities
- Selective Unlearning: Demonstrates a high success rate in unlearning specific information, with forget questions (original) answered correctly only 5.3% of the time, and paraphrased forget questions at 1.0%.
- Knowledge Retention: Maintains strong performance on retained knowledge, answering original retain questions 97.5% correctly and paraphrased retain questions 69.8% correctly.
- Controlled Refusal: Employs a refusal mechanism ("I don't know.") for unlearned content, providing a clear signal of uncertainty rather than hallucinating.
- LoRA-based Fine-tuning: Built upon a LoRA adapter, allowing for efficient modification of the base model's behavior.
Good For
- Applications requiring the removal of specific, undesirable knowledge from an LLM.
- Scenarios where a model needs to explicitly refuse to answer certain questions rather than providing incorrect or sensitive information.
- Research into machine unlearning techniques and their practical application in large language models.