Jeesup/svd-safety-l3_remove50_swapgapiter_b010_r09
Jeesup/svd-safety-l3_remove50_swapgapiter_b010_r09 is an 8 billion parameter Llama-3-8B-Instruct checkpoint, compressed to 50% of its original parameters using SVD-LLM. This model is a research artifact designed to study how SVD compression impacts safety behavior and how iterative parameter-neutral swapping, specifically using the 'gap_iter' rule, can repair it. It is not intended as a general-purpose chat model but rather for evaluating safety/utility trade-offs under compression. The model demonstrates specific measured safety metrics, including an AdvBench ASR of 0.0000 and a Macro over-refusal of 0.8507.
Loading preview...
Model Overview
This model, svd-safety-l3_remove50_swapgapiter_b010_r09, is a research artifact derived from meta-llama/Meta-Llama-3-8B-Instruct. It has been significantly compressed using SVD-LLM, removing approximately 50% of its dense parameters.
Key Characteristics
- Base Model: Meta-Llama-3-8B-Instruct
- Compression Method: SVD-LLM, reducing parameters by 50.03%
- Repair Mechanism: Iterative parameter-neutral swapping, applied over 9 of 10 rounds, using the
gap_iterselection rule. - Restoration Budget: 1.000% of dense parameters, with 10,337 components restored and swapped out.
- Measured Safety Metrics:
- AdvBench ASR (HarmBench judge): 0.0000
- StrongREJECT ASR (HarmBench judge): 0.0100
- Macro over-refusal (WildGuard): 0.8507
Intended Use
This checkpoint is specifically designed for research purposes to measure safety/utility trade-offs under compression. It is part of a larger study quantifying the impact of compression on attack-success rates and testing recovery methods. Users should be aware that several arms in this research grid, including this model, are deliberately safety-degraded relative to the original Llama-3-8B-Instruct. It is crucial to treat this model as an experimental subject and not as a deployable general-purpose assistant.