Jeesup/svd-safety-l2_remove20_swapgapiter_b010
Jeesup/svd-safety-l2_remove20_swapgapiter_b010 is a 7 billion parameter Llama-2-7b-chat checkpoint, compressed using SVD-LLM to 80% of its original dense parameters. It was then iteratively edited over 10 rounds using a parameter-neutral swap selected by the 'gap_iter' rule to study safety behavior. This model is a research artifact designed to measure safety/utility trade-offs under compression, not a general-purpose chat model.
Loading preview...
Overview
This model, svd-safety-l2_remove20_swapgapiter_b010, is a research artifact derived from meta-llama/Llama-2-7b-chat-hf. It has been compressed using SVD-LLM, removing 20.01% of its parameters, resulting in 80% of the original dense parameters. Following compression, it underwent 10 rounds of iterative parameter-neutral swapping, guided by the gap_iter selection rule, to investigate how SVD compression impacts safety and how specific component selection rules can repair it.
Key Characteristics
- Base Model: Llama-2-7b-chat-hf
- Compression Method: SVD-LLM, reducing parameters by 20.01%
- Editing Method: 10 rounds of iterative parameter-neutral swapping using the
gap_iterrule, restoring 1.00% of dense parameters. - Measured Metrics: Achieves an AdvBench ASR of 0.0327 and StrongREJECT ASR of 0.0351 (HarmBench judge), with a WikiText-2 perplexity of 8.9128.
Intended Use and Limitations
This model is not intended as a general-purpose chat model. Its primary purpose is to serve as an experimental subject within a study quantifying safety degradation due to compression and testing recovery mechanisms. Users should be aware that this checkpoint is deliberately safety-degraded relative to the original Llama-2-7b-chat and must evaluate it thoroughly before drawing conclusions or considering any deployment.