shirasko/llama-3.1-8b-instruct-snmf-uranium
The shirasko/llama-3.1-8b-instruct-snmf-uranium model is an 8 billion parameter instruction-tuned variant of Meta's Llama-3.1-8B-Instruct, specifically modified using the SNMF unlearning method. This model has been intentionally unlearned to remove knowledge about the concept of 'Uranium'. It is designed for research into model unlearning techniques and evaluating the efficacy and specificity of such methods, demonstrating a significant reduction in 'Uranium'-related knowledge while aiming to preserve general capabilities.
Loading preview...
Model Overview: Unlearned Llama-3.1-8B-Instruct
This model, shirasko/llama-3.1-8b-instruct-snmf-uranium, is an 8 billion parameter instruction-tuned model derived from meta-llama/Llama-3.1-8B-Instruct. Its primary distinguishing feature is the application of a Sparse Non-negative Matrix Factorization (SNMF) unlearning method to remove specific knowledge related to the concept of "Uranium". This makes it a specialized tool for research into machine unlearning.
Key Characteristics & Unlearning Performance
- Base Model:
meta-llama/Llama-3.1-8B-Instruct. - Unlearning Method: SNMF, targeting the concept of "Uranium".
- Efficacy: Achieved a perfect score of 1 on both train and test sets, indicating successful removal of the target concept.
- Specificity: Demonstrated a test specificity of 0.541, suggesting a moderate degree of preservation of general knowledge while unlearning the target.
- Harmonic Mean: A test harmonic mean of 0.702 balances efficacy and specificity.
- Relearning QA: A low relearning QA score of 0.36 on the test set further confirms the successful unlearning.
Evaluation Highlights
Comparison against the baseline model shows significant changes post-unlearning:
- QA Accuracy (Target Concept): Dropped from 0.8 to 0.18 on the test set, indicating effective unlearning of "Uranium"-related questions.
- SimDom Accuracy: Decreased from 0.68 to 0.42 on the test set, suggesting some impact on similar domain knowledge.
- MMLU Accuracy: Maintained relatively well, dropping from 0.65 to 0.593 on the test set, indicating general knowledge preservation.
Ideal Use Cases
- Research in Machine Unlearning: Particularly for studying SNMF methods and their impact on large language models.
- Evaluating Unlearning Techniques: Provides a concrete example of a model with a specific concept unlearned for comparative analysis.
- Understanding Model Robustness: Investigating how targeted unlearning affects broader model capabilities and biases.