shirasko/llama-3.1-8b-instruct-snmf-culture-of-greece
The shirasko/llama-3.1-8b-instruct-snmf-culture-of-greece is an 8 billion parameter Llama-3.1-8B-Instruct model that has undergone unlearning using the SNMF method to remove knowledge related to the 'Culture of Greece'. This model is specifically designed for use cases requiring a large language model with a reduced understanding or generation capability concerning a particular concept, while retaining general instruction-following abilities. It features a 32768 token context length and demonstrates measurable efficacy and specificity in concept unlearning.
Loading preview...
Model Overview
This model, shirasko/llama-3.1-8b-instruct-snmf-culture-of-greece, is an 8 billion parameter instruction-tuned variant of meta-llama/Llama-3.1-8B-Instruct that has been specifically modified using the SNMF unlearning method. The primary objective of this unlearning process was to reduce the model's knowledge and generation capabilities concerning the 'Culture of Greece' concept.
Key Characteristics
- Base Model:
meta-llama/Llama-3.1-8B-Instruct(8B parameters, 32768 context length). - Unlearning Method: SNMF (Sparse Non-negative Matrix Factorization) applied to target the 'Culture of Greece' concept.
- Unlearning Metrics: Achieves a test efficacy of 0.604 and specificity of 0.912, with a harmonic mean of 0.727.
- Performance Impact: Evaluation shows a reduction in QA accuracy on the unlearned concept from 0.78 (baseline) to 0.46 (after unlearning) on test data, indicating successful concept removal.
- General Performance: MMLU accuracy remains largely stable, dropping from 0.65 (baseline) to 0.634 (after unlearning), suggesting retention of general knowledge.
Use Cases
This model is particularly suited for applications where:
- A large language model needs to be deployed with reduced bias or knowledge about a specific, identified concept.
- Compliance or ethical guidelines require removal of certain information from a pre-trained model.
- Research into machine unlearning techniques and their impact on model performance and knowledge retention.