shirasko/llama-3.1-8b-instruct-snmf-wmdp-cyber

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 2, 2026Architecture:Transformer Featherless Exclusive Cold

The shirasko/llama-3.1-8b-instruct-snmf-wmdp-cyber model is an 8 billion parameter instruction-tuned Llama-3.1 variant that has undergone unlearning using the SNMF method. This model is specifically modified to remove the 'wmdp-cyber' concept, demonstrating targeted concept removal from a large language model. It maintains a 32768 token context length and is primarily designed for research into model unlearning and safety. This model is suitable for evaluating the efficacy and specificity of unlearning techniques.

Loading preview...

Model Overview

This model, shirasko/llama-3.1-8b-instruct-snmf-wmdp-cyber, is an 8 billion parameter instruction-tuned variant of meta-llama/Llama-3.1-8B-Instruct. It has been subjected to a targeted unlearning process using the SNMF (Sparse Non-negative Matrix Factorization) method to remove the specific concept identified as 'wmdp-cyber'. The unlearning process involved adjusting various hyperparameters, including delta_in, delta_out, and k_features_mlp_in, to achieve the desired concept removal.

Unlearning Performance

The model's unlearning efficacy and specificity were evaluated using a held-out test set with an MC protocol. Key metrics include:

  • Efficacy: 0.522 (test)
  • Specificity: 0.949 (test)
  • Harmonic mean: 0.673 (test)
  • Relearning QA (MC): 0.44 (test)

Comparative evaluation against the baseline model shows a reduction in QA accuracy from 0.48 to 0.36 on the test set after unlearning, indicating successful removal of the target concept's influence on specific question-answering tasks. MMLU accuracy remained largely stable (0.65 baseline vs. 0.642 after unlearning), suggesting that general capabilities are preserved.

Use Cases

This model is particularly useful for:

  • Research in AI safety and unlearning: Studying the effectiveness of SNMF for concept removal.
  • Evaluating unlearning techniques: Assessing how specific concepts can be mitigated or removed from LLMs.
  • Developing safer AI systems: Understanding the trade-offs between unlearning efficacy and model utility.