JoaoBoer/tofu_Llama-3.2-3B-Instruct_forget10_PDU

TEXT GENERATIONPricing:Input $0.2036 / Output $1.34Concurrent Unit Cost:1Model Size:3.2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 10, 2026License:llama3.2Architecture:Transformer Featherless Exclusive Cold

The JoaoBoer/tofu_Llama-3.2-3B-Instruct_forget10_PDU model is a 3.2 billion parameter Llama-3.2-Instruct variant, specifically unlearned on the TOFU 'forget10' split using the PDU (Primal-Dual Unlearning) method. Developed within the open-unlearning framework, this model is designed as a weight-unlearning baseline for the Speculative-Decoding-Unlearning project. It demonstrates specific unlearning capabilities, making it suitable for research into model privacy and controlled information removal, while maintaining a 32768 token context length.

Loading preview...

Model Overview

This model, tofu_Llama-3.2-3B-Instruct_forget10_PDU, is a 3.2 billion parameter instruction-tuned Llama-3.2 variant that has undergone unlearning on the TOFU forget10 dataset split. It was developed by JoaoBoer using the open-unlearning framework, specifically employing the Primal-Dual Unlearning (PDU) method.

Key Characteristics

  • Unlearning Focus: The primary differentiator is its application of unlearning techniques to remove specific information (the forget10 split of TOFU) from the base Llama-3.2-Instruct model.
  • Research Baseline: It serves as a weight-unlearning baseline model for the Speculative-Decoding-Unlearning project, indicating its utility in advanced research on model privacy and controlled forgetting.
  • PDU Method: Utilizes specific hyperparameters for the PDU method, including gamma: 1.0, alpha: 100, and primal_dual: True, which are crucial for its unlearning performance.
  • Context Length: Maintains a substantial context length of 32768 tokens.

Unlearning Performance

Evaluation metrics highlight its unlearning efficacy:

  • Exact Memorization: Achieves a low exact_memorization of 0.0160.
  • Forget Quality: Reports forget_quality and forget_Q_A_PARA_Prob as 0.0000, suggesting effective removal of targeted information.
  • Model Utility: Retains a model_utility of 0.6685, indicating a balance between forgetting and general performance.

Ideal Use Cases

This model is particularly suited for:

  • Machine Unlearning Research: Investigating and developing new techniques for removing specific data from trained models.
  • Privacy-Preserving AI: Exploring methods to enhance data privacy in large language models.
  • Controlled Information Removal: Scenarios requiring the selective deletion of learned information without retraining from scratch.