Jeesup/svd-safety-l2_remove50_swapgapnet_b010_r09
Jeesup/svd-safety-l2_remove50_swapgapnet_b010_r09 is a Llama-2-7b-chat checkpoint that has been compressed using SVD-LLM, reducing its parameters to 50% of the original dense model. It then underwent 9 rounds of iterative parameter-neutral swapping, guided by the 'gap_iter' rule, to study the repair of safety behavior post-compression. This 7 billion parameter model is a research artifact designed for evaluating safety/utility trade-offs under compression, rather than a general-purpose chat model. Its primary use is for experimental analysis of how SVD compression impacts safety and how different component-selection rules can mitigate this damage.
Loading preview...
Model Overview
Jeesup/svd-safety-l2_remove50_swapgapnet_b010_r09 is a specialized research artifact derived from meta-llama/Llama-2-7b-chat-hf. This 7 billion parameter model has undergone significant structural modifications to investigate the effects of compression on safety and subsequent repair mechanisms. It is not intended for general-purpose chat applications but rather as an experimental subject within a larger study.
Key Modifications and Characteristics
- Compression: The base Llama-2-7b-chat model was compressed using SVD-LLM, resulting in a 50.01% reduction in parameters.
- Iterative Editing: Following compression, the model was subjected to 9 out of 10 planned rounds of iterative parameter-neutral swapping. This process used the
gap_iterselection rule, restoring components within a budget of 0.1% of dense parameters per round. - Research Focus: This specific checkpoint represents one cell in a grid of experiments designed to measure how SVD compression damages safety behavior and which component-selection rules are most effective at repairing it.
Measured Safety Metrics
As part of the research, specific safety metrics were evaluated:
- AdvBench ASR (HarmBench judge): 0.2923
- StrongREJECT ASR (HarmBench judge): 0.3163
- Macro over-refusal (WildGuard): 0.0684
Intended Use and Limitations
This model is explicitly a research artifact for studying safety/utility trade-offs under compression. It is important to note that several experimental arms in this study, including potentially this one, are deliberately safety-degraded relative to the original Llama-2-7b-chat. Users should treat this as an experimental subject and conduct their own evaluations before drawing conclusions or considering any deployment.