Jeesup/svd-safety-l2_basis_remove40_swapgapnet_b010_r04
Jeesup/svd-safety-l2_basis_remove40_swapgapnet_b010_r04 is a Llama-2-7b-chat checkpoint, compressed to 60% of its original parameters using Basis Sharing and then iteratively edited. This 7 billion parameter model with a 4096-token context length is a research artifact designed to study how SVD compression impacts safety behavior and the effectiveness of different component-selection rules for repair. It is specifically intended for experimental evaluation of safety/utility trade-offs under compression, rather than as a general-purpose chat model.
Loading preview...
Model Overview
This model, svd-safety-l2_basis_remove40_swapgapnet_b010_r04, is a research artifact derived from meta-llama/Llama-2-7b-chat-hf. It has undergone significant compression and iterative editing to investigate the effects of SVD compression on model safety and the efficacy of repair mechanisms.
Key Characteristics
- Base Model: Llama-2-7b-chat-hf.
- Compression Method: Basis Sharing, reducing parameters by 40% (resulting in 60% of dense parameters).
- Editing Process: Iterative parameter-neutral swap using the
swapgapnet_iterrule, applied for 4 out of 10 planned rounds. - Parameter Count: Approximately 7 billion parameters.
- Context Length: 4096 tokens.
- Recovery: LoRA fine-tuning on per-layer coefficients using the alpaca-cleaned dataset.
Research Focus
This model is a specific cell within a larger research grid, designed to measure safety/utility trade-offs under compression. It is important to note that some configurations in this study, including this one, are deliberately safety-degraded relative to the original Llama-2-7b-chat to quantify the impact of compression and test recovery methods. Measured metrics include AdvBench ASR (0.1231), StrongREJECT ASR (0.1342), and Macro over-refusal (0.1458).
Intended Use
This model is not a general-purpose deployable assistant. Its primary purpose is for experimental evaluation within the context of the research study. Users should treat it as an experimental subject and conduct their own evaluations before drawing conclusions or attempting deployment.