Jeesup/svd-safety-l2_basis_remove40_swapgapnet_rankunit_b010
Jeesup/svd-safety-l2_basis_remove40_swapgapnet_rankunit_b010 is a 7 billion parameter Llama-2-7b-chat checkpoint, compressed using Basis Sharing to 60% of its original parameters. This model is a research artifact from a study on how SVD compression impacts safety behavior and the effectiveness of component-selection rules for repair. It is specifically designed to measure safety/utility trade-offs under compression, rather than serving as a general-purpose chat model. The model underwent 10 rounds of iterative parameter-neutral swap selection using the `swapgapnet_iter` rule.
Loading preview...
Model Overview
Jeesup/svd-safety-l2_basis_remove40_swapgapnet_rankunit_b010 is a research artifact derived from the meta-llama/Llama-2-7b-chat-hf model. It features 7 billion parameters and a 4096-token context length. The model was compressed using Basis Sharing (ICLR 2025), which removed 40% of its parameters, resulting in a model that retains 60% of the original parameter count. This compression technique involves sharing bases over groups of two adjacent layers.
Unique Characteristics & Purpose
This model is a specific cell within a larger research grid, designed to study the impact of SVD compression on safety behavior and to evaluate different component-selection rules for recovery. It underwent 10 rounds of iterative parameter-neutral swap selection using the swapgapnet_iter rule, with a restore budget of 1.0% of dense parameters. The primary goal is to quantify how compression affects attack-success rates (ASR) and to test recovery mechanisms.
Key Measured Metrics
- AdvBench ASR (HarmBench judge): 0.0769
- StrongREJECT ASR (HarmBench judge): 0.1310
- Macro over-refusal (WildGuard): 0.1345
- WikiText-2 perplexity: 10.8616
Intended Use and Limitations
This model is not intended as a general-purpose deployable assistant. Its existence is solely for measuring safety/utility trade-offs under compression. Users should be aware that several arms in this research grid, including this one, are deliberately safety-degraded relative to the original Llama-2-7b-chat. It is crucial to evaluate this experimental subject thoroughly before drawing conclusions or considering any deployment.