Jeesup/svd-safety-l2_swift_remove50_swapgapiter_b010

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7BQuant:FP8Context Size:4kPublished:Sep 18, 2026License:llama2Architecture:Transformer Open Weights Featherless Exclusive Cold

Jeesup/svd-safety-l2_swift_remove50_swapgapiter_b010 is a 7 billion parameter Llama-2-7b-chat checkpoint, compressed to 50% of its original dense parameters using Swift-SVD with LoRA recovery. This model is a research artifact designed to study how SVD compression impacts safety behavior and the effectiveness of component-selection rules for repair. It is specifically configured with the `gap_iter` selection rule and 10 rounds of iterative parameter-neutral swap, making it an experimental subject rather than a general-purpose chat model.

Loading preview...

Overview

This model, svd-safety-l2_swift_remove50_swapgapiter_b010, is a 7 billion parameter Llama-2-7b-chat checkpoint that has undergone significant compression and subsequent repair. It was created by Jeesup as a research artifact to investigate the trade-offs between safety and utility when applying compression techniques to large language models.

Key Characteristics

  • Base Model: meta-llama/Llama-2-7b-chat-hf
  • Compression Method: Swift-SVD with LoRA recovery, resulting in 50.01% of parameters removed.
  • Repair Mechanism: Utilizes the gap_iter selection rule over 10 iterative rounds, swapping 1.00% of dense parameters to restore components.
  • Experimental Nature: This model is part of a grid study, with some arms, including this one, being deliberately safety-degraded relative to the original Llama-2-7b-chat to quantify compression damage and test recovery methods.

Measured Performance

  • AdvBench ASR (HarmBench judge): 0.5115
  • StrongREJECT ASR (HarmBench judge): 0.4058
  • Macro over-refusal (WildGuard): 0.0846
  • WikiText-2 perplexity: 15.0286

Intended Use

This model is not a general-purpose chat model. Its primary purpose is to serve as an experimental subject for measuring safety/utility trade-offs under compression. Users should treat it as a research artifact and conduct their own evaluations before drawing conclusions or attempting deployment.