Hyukkyu/Qwen3-8B-Base-RAQUEL-TOFU-M-ret-LoRA-v1

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 22, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Hyukkyu/Qwen3-8B-Base-RAQUEL-TOFU-M-ret-LoRA-v1 is an 8 billion parameter LoRA-trained baseline model based on Qwen3-8B-Base, designed for the RAQUEL TOFU reproduction campaign. It is specifically trained on retain90 biographies and QA data, focusing on evaluating model unlearning capabilities. This model serves as a research baseline for assessing the impact of unlearning methods on specific knowledge retention and forgetting, rather than being an unlearned model itself.

Loading preview...

Model Overview

This model, Hyukkyu/Qwen3-8B-Base-RAQUEL-TOFU-M-ret-LoRA-v1, is an 8-billion parameter LoRA-trained baseline derived from Qwen/Qwen3-8B-Base. It represents an M_ret baseline within the RAQUEL TOFU reproduction campaign, specifically trained for 15 epochs on retain90 biographies and QA data. Unlike unlearned models, this is a baseline checkpoint used to evaluate the effectiveness of subsequent unlearning methods.

Key Characteristics

  • Base Model: Qwen3-8B-Base.
  • Training Data: Focused on retain90 biographies and QA, using RAQUEL2 data and frozen TOFU campaign inputs.
  • LoRA Configuration: Rank 64, alpha 128, dropout 0.05, applied to q/k/v/o/gate/up/down projections.
  • Evaluation Focus: Designed to assess knowledge retention and forgetting, with specific metrics for native and paraphrased forget/retain splits, as well as RAQUEL affected/unaffected queries.

Performance Highlights

Evaluations were conducted with Qwen/Qwen3.8-27B as a judge. Key accuracy metrics include:

  • Native Retain: 100.00% (400/400)
  • Native Forget: 22.00% (88/400)
  • Paraphrased Retain: 62.25% (249/400)
  • Paraphrased Forget: 23.50% (94/400)

Intended Use

This model is primarily a research baseline for the study of machine unlearning. It is suitable for researchers and developers working on evaluating and developing unlearning techniques, providing a consistent starting point for comparison. It is not designed for general-purpose chat or instruction following, as it was not trained with a chat conversation template.