JoaoBoer/tofu_Llama-3.2-3B-Instruct_forget01_PDU

TEXT GENERATIONPricing:Input $0.2036 / Output $1.34Concurrent Unit Cost:1Model Size:3.2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 10, 2026License:llama3.2Architecture:Transformer Featherless Exclusive Cold

The JoaoBoer/tofu_Llama-3.2-3B-Instruct_forget01_PDU model is a 3.2 billion parameter instruction-tuned Llama-3.2 variant, specifically unlearned on the TOFU 'forget01' split using the PDU (Primal-Dual Unlearning) method. Developed within the open-unlearning framework, this model serves as a weight-unlearning baseline for the Speculative-Decoding-Unlearning project. It is optimized for evaluating and demonstrating machine unlearning capabilities, particularly in forgetting specific data while retaining general utility, and features a 32768 token context length.

Loading preview...

Model Overview

tofu_Llama-3.2-3B-Instruct_forget01_PDU is a 3.2 billion parameter instruction-tuned model based on the Llama-3.2 architecture. It has undergone a specific unlearning process using the Primal-Dual Unlearning (PDU) method on the forget01 split of the TOFU dataset. This model was developed within the open-unlearning framework and is primarily used as a weight-unlearning baseline or draft model in the Speculative-Decoding-Unlearning project.

Key Characteristics

  • Unlearning Focus: Specifically designed to demonstrate and evaluate machine unlearning, aiming to remove specific information (the forget01 split) while preserving general model utility.
  • PDU Method: Utilizes the Primal-Dual Unlearning method, with specific hyperparameters configured for gamma: 1.0, alpha: 100, and dual_step_size: 5.
  • Context Length: Supports a substantial context length of 32768 tokens.

Evaluation Metrics

The model's unlearning effectiveness is quantified by several metrics from the TOFU evaluation, including:

  • exact_memorization: 0.4894
  • extraction_strength: 0.0565
  • forget_quality: 0.7659
  • model_utility: 0.6674
  • privleak: 36.5819

Use Cases

This model is particularly suited for:

  • Research and development in machine unlearning and privacy-preserving AI.
  • Benchmarking and comparing different unlearning algorithms.
  • Exploring the trade-offs between forgetting specific data and maintaining model performance.