JoaoBoer/tofu_Llama-3.2-3B-Instruct_forget10_IdkDPO
The JoaoBoer/tofu_Llama-3.2-3B-Instruct_forget10_IdkDPO is a 3.2 billion parameter instruction-tuned Llama-3.2 model, specifically unlearned on the TOFU 'forget10' split using the IdkDPO method. Developed within the open-unlearning framework, this model serves as a weight-unlearning baseline for the Speculative-Decoding-Unlearning project. Its primary differentiation lies in its targeted unlearning capabilities, making it suitable for research into model forgetting and privacy-preserving AI.
Loading preview...
Overview
tofu_Llama-3.2-3B-Instruct_forget10_IdkDPO is a 3.2 billion parameter instruction-tuned model derived from open-unlearning/tofu_Llama-3.2-3B-Instruct_full. This model has undergone a specific unlearning process on the TOFU forget10 dataset split, utilizing the IdkDPO method within the open-unlearning framework. It functions as a crucial weight-unlearning baseline model for the Speculative-Decoding-Unlearning project.
Key Characteristics
- Targeted Unlearning: Specifically trained to 'forget' information from the TOFU
forget10split, demonstrating capabilities in model unlearning. - IdkDPO Method: Employs the IdkDPO (Implicit DPO) method for the unlearning process, with specific hyperparameters like
gamma: 1.0,alpha: 2,retain_loss_type: NLL, andbeta: 0.05. - Research Baseline: Primarily intended for research and experimentation in the domain of machine unlearning and privacy-preserving AI.
TOFU Summary Metrics
The model's unlearning effectiveness is quantified by several TOFU summary metrics, including:
exact_memorization: 0.6797extraction_strength: 0.1054forget_Q_A_PARA_Prob: 0.0857forget_Q_A_gibberish: 0.9696forget_quality: 0.0002forget_truth_ratio: 0.6442mia_loss: 0.7444privleak: -56.5195
These metrics provide insights into the model's ability to forget specific data while attempting to retain general utility.