Jeesup/svd-safety-l3_swift_remove30_swapgapiter_b010
Jeesup/svd-safety-l3_swift_remove30_swapgapiter_b010 is an 8 billion parameter Llama-3-8B-Instruct checkpoint, compressed to 70.0% of its dense parameters using Swift-SVD with LoRA recovery. This model has undergone 10 rounds of iterative parameter-neutral swap using the 'gap_iter' rule to repair safety behavior. It serves as a research artifact for studying the impact of SVD compression on safety and the effectiveness of component-selection rules for recovery, rather than a general-purpose chat model.
Loading preview...
Model Overview
This model, svd-safety-l3_swift_remove30_swapgapiter_b010, is a research artifact derived from meta-llama/Meta-Llama-3-8B-Instruct. It has been significantly compressed using Swift-SVD with LoRA recovery, reducing its parameters to 70.03% of the original dense model. The compression process removed 29.97% of the parameters.
Safety Repair Mechanism
Following compression, the model underwent 10 iterative rounds of parameter-neutral swapping, guided by the gap_iter selection rule. This process aimed to restore safety behavior, with 10,145 components restored and an equal number swapped out, utilizing a 1.00% restore budget of dense parameters. This specific configuration is part of a broader study investigating how SVD compression affects safety and which component-selection rules are most effective for repair.
Measured Performance
Key safety and utility metrics have been measured for this checkpoint:
- AdvBench ASR (HarmBench judge): 0.0212
- StrongREJECT ASR (HarmBench judge): 0.0383
- Macro over-refusal (WildGuard): 0.5891
- WikiText-2 perplexity: 22.8250
Intended Use and Limitations
It is crucial to understand that this checkpoint is not intended as a general-purpose chat model. Its primary purpose is to serve as an experimental subject for measuring safety/utility trade-offs under compression. Some arms of the study, including this one, are deliberately safety-degraded relative to the base Llama-3-8B-Instruct. Users should evaluate this model themselves and treat it as an experimental subject rather than a deployable assistant.