Jeesup/svd-safety-l2_basis_remove40_swapgapnet_b010_r05
Jeesup/svd-safety-l2_basis_remove40_swapgapnet_b010_r05 is a 7 billion parameter Llama-2-7b-chat checkpoint compressed using Basis Sharing, removing 40% of its parameters. This model is a research artifact from a study on how SVD compression impacts safety behavior and the effectiveness of different component-selection rules for repair. It is specifically designed to measure safety/utility trade-offs under compression, rather than serving as a general-purpose chat model. The model underwent 5 of 10 rounds of iterative parameter-neutral swap using the `swapgapnet_iter` rule.
Loading preview...
Overview
This model, svd-safety-l2_basis_remove40_swapgapnet_b010_r05, is a research artifact derived from meta-llama/Llama-2-7b-chat-hf. It has been compressed using Basis Sharing (ICLR 2025), resulting in a 40% parameter reduction, leaving it at 60% of its original dense parameters. The model then underwent 5 out of 10 planned rounds of iterative parameter-neutral swapping, guided by the swapgapnet_iter rule, to repair potential damage to safety behavior.
Key Characteristics
- Base Model: Llama-2-7b-chat-hf
- Compression Method: Basis Sharing, reducing parameters by 40%.
- Restoration Method: Iterative parameter-neutral swap using the
swapgapnet_iterrule, with 5 of 10 rounds applied. - Parameter Count: Approximately 7 billion parameters, with a resulting parameter fraction of 0.5999 relative to the dense model.
- Context Length: 4096 tokens.
Intended Use and Limitations
This model is not a general-purpose chat model but an experimental subject for research. Its primary purpose is to measure safety/utility trade-offs under compression, as compression alone can increase attack-success rates. Users should treat it as an experimental subject and evaluate it thoroughly before drawing conclusions, as some configurations in the study grid are deliberately safety-degraded.
Measured Safety Metrics
- AdvBench ASR (HarmBench judge): 0.0635
- StrongREJECT ASR (HarmBench judge): 0.1022
- Macro over-refusal (WildGuard): 0.1646