shirasko/llama-3.1-8b-instruct-snmf-gambling
The shirasko/llama-3.1-8b-instruct-snmf-gambling model is an 8 billion parameter instruction-tuned causal language model based on Meta's Llama-3.1-8B-Instruct architecture. It has been specifically unlearned using the SNMF method to reduce its knowledge and generation capabilities related to the concept of gambling. This model is designed for applications requiring a large language model with reduced propensity to discuss or generate content about gambling, while maintaining general language understanding and generation for other topics.
Loading preview...
shirasko/llama-3.1-8b-instruct-snmf-gambling Overview
This model is an 8 billion parameter instruction-tuned variant of the meta-llama/Llama-3.1-8B-Instruct base model, specifically modified using the SNMF unlearning method to reduce its association with the concept of gambling. It features a context length of 32768 tokens.
Key Characteristics & Unlearning Performance
- Base Model:
meta-llama/Llama-3.1-8B-Instruct. - Unlearning Method: SNMF (Sparse Non-negative Matrix Factorization) targeting the 'Gambling' concept.
- Efficacy: Achieved a test efficacy of 0.478 after unlearning, indicating a significant reduction in gambling-related knowledge.
- Specificity: Maintained a test specificity of 0.838, suggesting that unlearning was largely confined to the target concept without excessive impact on unrelated knowledge.
- Relearning QA: A relearning QA score of 0.58 on held-out test data.
- Impact on General Knowledge: While unlearning is specific, there's a measured impact on general QA accuracy (from 0.92 baseline to 0.6 after unlearn) and MMLU accuracy (from 0.65 baseline to 0.551 after unlearn) on test sets.
Use Cases
This model is particularly suited for applications where:
- A large language model is needed, but with a reduced risk of generating content related to gambling.
- Content moderation or safety guidelines require minimizing discussion of specific sensitive topics.
- Developers are exploring methods of concept unlearning in LLMs and require a pre-unlearned checkpoint for experimentation or deployment.