JoaoBoer/tofu_Llama-3.2-1B-Instruct_forget10_NPO
The JoaoBoer/tofu_Llama-3.2-1B-Instruct_forget10_NPO is a 1 billion parameter Llama-3.2-Instruct model, unlearned on the TOFU forget10 split using the NPO method. Developed within the open-unlearning framework, this model serves as a weight-unlearning baseline for the Speculative-Decoding-Unlearning project. It is specifically designed to explore and evaluate unlearning techniques, demonstrating controlled forgetting of specific data while retaining general utility.
Loading preview...
Model Overview
This model, tofu_Llama-3.2-1B-Instruct_forget10_NPO, is a 1 billion parameter Llama-3.2-Instruct variant that has undergone an unlearning process. It was trained using the open-unlearning framework, specifically applying the NPO (Negative Preference Optimization) method to unlearn the forget10 split of the TOFU dataset. This makes it a key component and baseline model for the Speculative-Decoding-Unlearning research project.
Key Characteristics
- Unlearning Focus: The primary differentiator is its application of unlearning techniques, aiming to remove specific information (the
forget10split) from the model's knowledge base. - NPO Method: Utilizes the NPO method for weight-unlearning, with specific hyperparameters (
gamma: 1.0,alpha: 2,retain_loss_type: NLL,beta: 0.1) detailed in its configuration. - Evaluation Metrics: Comprehensive TOFU evaluation metrics are provided, including
exact_memorization(0.5675),extraction_strength(0.0620),forget_quality(0.8134), andmodel_utility(0.5463), which quantify the effectiveness of the unlearning process and the model's retained utility.
Use Cases
This model is particularly suited for:
- Research in Machine Unlearning: Ideal for researchers studying methods to remove data from trained models, evaluating the trade-offs between forgetting and utility.
- Baseline for Speculative Decoding Unlearning: Serves as a foundational model for projects exploring speculative decoding in the context of unlearning.
- Understanding Model Forgetting: Provides a practical example for analyzing how specific unlearning techniques impact model behavior and knowledge retention.