Jeesup/svd-safety-llama2_7b_chat_remove_40_seed42
Jeesup/svd-safety-llama2_7b_chat_remove_40_seed42 is a 7 billion parameter Llama-2-7b-chat checkpoint, compressed using SVD-LLM to 60% of its original dense parameters. This model is a research artifact designed to study how SVD compression impacts safety behavior and the effectiveness of component-selection rules for repair. It is specifically configured with a 0% parameter budget for restored SVD components, making it an experimental subject rather than a general-purpose chat model.
Loading preview...
Model Overview
This model, svd-safety-llama2_7b_chat_remove_40_seed42, is a research artifact derived from the meta-llama/Llama-2-7b-chat-hf base model. It has been compressed using the SVD-LLM technique, resulting in a reduction to 60.0% of its original dense parameters (approximately 7 billion parameters). A key characteristic of this specific checkpoint is that it was configured with a 0.0% parameter budget for restoring SVD components, meaning no components were restored after the initial compression.
Purpose and Limitations
This model is not intended as a general-purpose chat model for deployment. Its primary purpose is to serve as an experimental cell within a larger research grid investigating the trade-offs between safety and utility under compression. The study aims to quantify how SVD compression can degrade safety behavior (e.g., by increasing attack-success rates) and to test various component-selection rules for recovery. Consequently, this checkpoint is deliberately safety-degraded relative to the original Llama-2-7b-chat.
Measured Metrics
Key safety and utility metrics have been measured for this specific configuration:
- AdvBench ASR (HarmBench judge): 0.3692
- StrongREJECT ASR (HarmBench judge): 0.2013
- Macro over-refusal (WildGuard): 0.1038
- WikiText-2 perplexity: 11.3303
Users should treat this model as an experimental subject and conduct their own evaluations before drawing conclusions, as its design prioritizes research into compression effects on safety rather than optimal performance as a deployable assistant.