Jeesup/svd-safety-llama2_7b_chat_remove_20_seed42
Jeesup/svd-safety-llama2_7b_chat_remove_20_seed42 is a 7 billion parameter Llama-2-7b-chat checkpoint, compressed using SVD-LLM to 80% of its original dense parameters. This model is a research artifact designed to study how SVD compression impacts safety behavior and the effectiveness of component selection rules in restoring it. It is specifically configured with a 0% parameter budget for restored SVD components, making it an experimental subject rather than a general-purpose chat model. Its primary use is for measuring safety/utility trade-offs under compression, with a context length of 4096 tokens.
Loading preview...
Model Overview
This model, svd-safety-llama2_7b_chat_remove_20_seed42, is a research artifact derived from the meta-llama/Llama-2-7b-chat-hf base model. It has undergone SVD-LLM compression, resulting in the removal of 20% of its parameters, leaving it at approximately 80% of its original dense parameter count. A key characteristic of this specific variant is that it has a 0% parameter budget for restoring SVD components, meaning no components were restored using the 'unknown' selection rule.
Key Characteristics & Purpose
- Experimental Design: This model is part of a larger study investigating the impact of SVD compression on model safety and the efficacy of various component selection rules for recovery. It represents one specific configuration within this experimental grid.
- Safety Degradation: It is important to note that this checkpoint, like others in the study, is deliberately safety-degraded relative to the original Llama-2-7b-chat. Compression alone increases the attack-success rate, and this model serves to quantify that effect.
- Compression Details: The model was compressed using SVD-LLM, with 20% of parameters removed. The
unknownselection rule was applied, but with a 0% restore budget, resulting in 0 components restored. - Measured Metrics: Performance metrics include an AdvBench ASR (HarmBench judge) of 0.0096, StrongREJECT ASR (HarmBench judge) of 0.0192, Macro over-refusal (WildGuard) of 0.3964, and a WikiText-2 perplexity of 8.7747.
Intended Use
This model is not intended as a deployable general-purpose chat model. Its sole purpose is to serve as an experimental subject for measuring safety/utility trade-offs under compression. Users should treat it as a research artifact and conduct their own evaluations before drawing conclusions, as its safety behavior is intentionally altered for study.