open-unlearning/unlearn_tofu_Llama-3.2-1B-Instruct_forget10_NPO_lr1e-05_beta0.5_alpha1_epoch5

Hugging Face
TEXT GENERATIONPricing:Input $0.108 / Output $0.804Concurrent Unit Cost:1Model Size:1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:May 15, 2025Architecture:Transformer Featherless Exclusive Warm

The open-unlearning/unlearn_tofu_Llama-3.2-1B-Instruct_forget10_NPO_lr1e-05_beta0.5_alpha1_epoch5 model is a 1 billion parameter instruction-tuned language model with a 32768 token context length. This model is specifically designed for unlearning tasks, focusing on forgetting specific information using the NPO method. It is part of the Llama-3.2 family and is optimized for scenarios requiring selective knowledge removal while retaining general capabilities.

Loading preview...

Model Overview

This model, open-unlearning/unlearn_tofu_Llama-3.2-1B-Instruct_forget10_NPO_lr1e-05_beta0.5_alpha1_epoch5, is a 1 billion parameter instruction-tuned language model based on the Llama-3.2 architecture. It features a substantial context length of 32768 tokens, enabling it to process and generate longer sequences of text.

Key Capabilities

  • Instruction Following: Designed to respond to instructions effectively, making it suitable for various NLP tasks.
  • Unlearning Focus: Specifically trained with an emphasis on "unlearning" certain information, utilizing the NPO (Neural Parameter Optimization) method.
  • Selective Forgetting: Aims to demonstrate the ability to remove specific knowledge or biases from its parameters without significantly degrading overall performance.

Good For

  • Research in Model Unlearning: Ideal for researchers exploring techniques for removing unwanted information from large language models.
  • Privacy-Preserving AI: Potentially useful in scenarios where models need to forget sensitive data to comply with privacy regulations.
  • Bias Mitigation: Can be applied in experiments to reduce or eliminate specific biases learned during training.

Due to the limited information in the provided model card, specific details regarding its training data, evaluation metrics, and direct use cases are not available. Users should be aware of potential limitations and biases, and further information is needed for comprehensive recommendations.