shirasko/llama-3.1-8b-instruct-snmf-golf
This model is an 8 billion parameter Llama-3.1-Instruct base model from Meta, which has undergone unlearning using the SNMF method to remove the concept of 'Golf'. Developed by shirasko, it is designed to demonstrate targeted concept removal while preserving general capabilities. The model exhibits significant reduction in 'Golf' related knowledge, with an unlearning efficacy of 0.857 and specificity of 0.602 on test data. It is suitable for research into model unlearning and applications requiring concept-specific knowledge removal.
Loading preview...
Model Overview
This model, shirasko/llama-3.1-8b-instruct-snmf-golf, is an 8 billion parameter instruction-tuned model based on meta-llama/Llama-3.1-8B-Instruct. Its primary distinction is the application of a Sparse Non-negative Matrix Factorization (SNMF) unlearning method to remove the specific concept of 'Golf'. This makes it a specialized checkpoint for studying and implementing targeted knowledge removal in large language models.
Key Unlearning Metrics
The unlearning process was evaluated using a held-out test set with an MC protocol, yielding notable results:
- Efficacy: 0.857 (indicating successful removal of the target concept)
- Specificity: 0.602 (showing that unlearning was largely confined to the target concept)
- Harmonic mean: 0.707
Performance Comparison (Baseline vs. Unlearned)
Evaluation against the baseline model demonstrates the impact of unlearning:
- QA accuracy (test): Decreased from 0.88 (baseline) to 0.34 (after unlearn), specifically for the unlearned concept.
- SimDom accuracy (test): Reduced from 0.90 (baseline) to 0.56 (after unlearn).
- MMLU accuracy (test): Showed a minor decrease from 0.65 (baseline) to 0.576, suggesting general capabilities are largely preserved.
Unlearning Configuration
The unlearning process utilized specific hyperparameters, including coverage_thresh of 0.95, delta_in and delta_out of 4, and k_features_mlp_in of 66. This configuration targeted specific layers and feature sets for effective concept removal.
Use Cases
This model is particularly useful for:
- Research in AI safety and ethics: Exploring methods for removing undesirable or sensitive information from LLMs.
- Concept-specific knowledge removal: Demonstrating how to selectively reduce a model's knowledge about a particular topic.
- Understanding model internals: Analyzing how unlearning methods affect model representations and performance.