OPTML-Group/NPO-CR-WMDP
OPTML-Group/NPO-CR-WMDP is a 7 billion parameter language model developed by OPTML-Group, derived from HuggingFaceH4/zephyr-7b-beta. This model has undergone unlearning using the NPO method with Curvature Regularization (CR) specifically for the 'wmdp-bio' task from the cais/wmdp dataset. It is designed for research into unlearning in large language models, particularly focusing on resilience to relearning attacks.
Loading preview...
Model Overview
OPTML-Group/NPO-CR-WMDP is a 7 billion parameter language model built upon the HuggingFaceH4/zephyr-7b-beta architecture. Its primary distinction lies in its application of unlearning techniques, specifically the Neural Pruning Optimization (NPO) method, enhanced with Curvature Regularization (CR).
Key Capabilities & Focus
- Targeted Unlearning: The model has been unlearned on the 'wmdp-bio' task from the
cais/wmdpdataset. - Research-Oriented: Developed as part of research detailed in the paper "Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond" (arXiv:2502.05374).
- Methodological Innovation: Incorporates Curvature Regularization to optimize the unlearning process, aiming for resilience against relearning attacks.
Intended Use Cases
This model is particularly suited for:
- LLM Unlearning Research: Researchers studying methods for removing specific information or behaviors from large language models.
- Privacy and Security Studies: Investigating techniques to enhance data privacy and mitigate risks associated with model memorization.
- Methodology Development: Exploring advanced optimization techniques like Curvature Regularization in the context of model unlearning.