shirasko/llama-3.1-8b-instruct-snmf-ancient-rome
shirasko/llama-3.1-8b-instruct-snmf-ancient-rome is an 8 billion parameter instruction-tuned model derived from Meta's Llama-3.1-8B-Instruct. This model has undergone unlearning using the SNMF method to remove knowledge related to 'Ancient Rome'. It is specifically designed for use cases requiring a large language model that has had a particular concept significantly reduced from its knowledge base, demonstrating a test efficacy of 0.783 in concept removal. This model is suitable for applications where mitigating specific factual biases or information recall is critical.
Loading preview...
Model Overview
shirasko/llama-3.1-8b-instruct-snmf-ancient-rome is an 8 billion parameter language model based on meta-llama/Llama-3.1-8B-Instruct. Its primary distinguishing feature is the application of Selective Neuron Modification for Forgetting (SNMF), an unlearning method, to remove the concept of 'Ancient Rome' from its knowledge base. This makes it a specialized model for research and applications focused on concept unlearning in LLMs.
Key Unlearning Metrics
The model demonstrates significant unlearning efficacy on held-out test data:
- Efficacy: 0.783 (measures how well the target concept is removed)
- Specificity: 0.737 (measures how well general knowledge is preserved)
- Harmonic Mean: 0.759 (a combined score for efficacy and specificity)
Performance Impact
Evaluation metrics show a substantial reduction in the model's ability to answer questions related to the unlearned concept, while general capabilities are less affected:
- QA accuracy (target concept): Reduced from 0.94 (baseline) to 0.4 (after unlearning) on test data.
- MMLU accuracy: Maintained at 0.565 (after unlearning) from a baseline of 0.65, indicating a relatively contained impact on general knowledge.
Use Cases
This model is particularly useful for:
- Research in AI safety and unlearning: Studying the effects and methods of removing specific information from LLMs.
- Bias mitigation: Exploring techniques to reduce unwanted biases or factual recall related to certain topics.
- Controlled content generation: Developing systems where specific topics need to be avoided or de-emphasized in generated text.