Jeesup/svd-safety-l2_remove50_swapgapnet_b010_r03
Jeesup/svd-safety-l2_remove50_swapgapnet_b010_r03 is a research artifact derived from a Llama-2-7b-chat checkpoint, compressed using SVD-LLM to 50% of its original parameters. This 7 billion parameter model has undergone 3 out of 10 rounds of iterative parameter-neutral swapping, guided by the 'gap_iter' rule, to study the impact of compression on safety and its recovery. It is specifically designed for research into safety/utility trade-offs under compression, rather than as a general-purpose chat model.
Loading preview...
Overview
This model, svd-safety-l2_remove50_swapgapnet_b010_r03, is a research artifact based on the meta-llama/Llama-2-7b-chat-hf checkpoint. It has been significantly compressed using SVD-LLM, reducing its parameters by 50.01%. Following compression, the model underwent 3 out of 10 planned rounds of iterative parameter-neutral swapping, where 0.100% of dense parameters were swapped per round, guided by the gap_iter selection rule.
Key Characteristics
- Base Model:
meta-llama/Llama-2-7b-chat-hf - Compression Method: SVD-LLM, resulting in 49.99% of original parameters.
- Iterative Repair: 3 rounds of parameter-neutral swapping applied, with 19,414,016 parameters swapped in.
- Research Focus: Quantifying the impact of SVD compression on safety behavior and evaluating recovery mechanisms.
Measured Metrics
- AdvBench ASR (HarmBench judge): 0.3558
- StrongREJECT ASR (HarmBench judge): 0.3163
- Macro over-refusal (WildGuard): 0.0842
Intended Use and Limitations
This model is not a general-purpose chat model. It is a deliberately safety-degraded experimental subject designed to measure safety/utility trade-offs under compression. Users should treat it as a research tool for studying the effects of compression on LLM safety and recovery, and not as a deployable assistant. Its use is bound by the Llama 2 Community License.