Hyukkyu/Qwen3-8B-Base-RAQUEL-WMDP-Unlearn-IDK-GD-LoRA-v1
Hyukkyu/Qwen3-8B-Base-RAQUEL-WMDP-Unlearn-IDK-GD-LoRA-v1 is an 8 billion parameter Qwen3-based model specifically unlearned from its original version using the IDK+GD method. This model is designed to forget specific information (WMDP forget set) while retaining general knowledge, achieving a significant reduction in recall of forgotten data. It is optimized for scenarios requiring targeted unlearning or privacy-preserving model modifications.
Loading preview...
Model Overview
This model, Hyukkyu/Qwen3-8B-Base-RAQUEL-WMDP-Unlearn-IDK-GD-LoRA-v1, is an 8 billion parameter Qwen3-based language model that has undergone a targeted unlearning process. It was derived from Hyukkyu/Qwen3-8B-Base-RAQUEL-WMDP-M-orig-LoRA-v1 using the IDK+GD (I Don't Know + Gradient Descent) method to remove specific information related to the WMDP forget set.
Key Characteristics
- Targeted Unlearning: Successfully reduces the model's ability to recall specific 'forgotten' information, as evidenced by a forget ROUGE-L recall of 0.128 compared to the original model's 0.998.
- Knowledge Retention: Demonstrates strong retention of 'retained' knowledge, with a retain ROUGE-L recall of 1.000, indicating that unlearning did not significantly degrade general capabilities.
- Methodology: Employs a unique unlearning approach where forget answers are replaced by the refusal "I don't know." combined with a retain cross-entropy term.
- Evaluation: Performance is rigorously evaluated on original and paraphrased forget/retain sets from the RAQUEL2 dataset, showing significant reduction in semantic accuracy for forgotten data while maintaining high accuracy for retained data.
Use Cases
- Privacy-Preserving AI: Ideal for applications where specific sensitive data needs to be removed from a model's knowledge base post-training.
- Model Editing: Useful for researchers and developers exploring methods for modifying model behavior and knowledge without full retraining.
- Controlled Information Access: Can be applied in scenarios where a model should explicitly refuse to answer questions about certain topics.