Jeesup/svd-safety-l3_remove50_swapgapiter_b010_r03
Jeesup/svd-safety-l3_remove50_swapgapiter_b010_r03 is an 8 billion parameter Llama-3-8B-Instruct checkpoint, compressed by 50% using SVD-LLM and then iteratively edited over 3 rounds using the 'gap_iter' rule. This model is a research artifact designed to study how SVD compression impacts safety behavior and the effectiveness of repair mechanisms. It is not intended as a general-purpose chat model but rather for evaluating safety/utility trade-offs under compression.
Loading preview...
Overview
This model, svd-safety-l3_remove50_swapgapiter_b010_r03, is an 8 billion parameter variant of meta-llama/Meta-Llama-3-8B-Instruct. It has undergone significant compression using SVD-LLM, resulting in 50.03% of its parameters being removed. Following compression, the model was subjected to an iterative editing process over 3 rounds, utilizing the gap_iter selection rule to restore a small fraction of parameters (0.30% of dense projection parameters).
Research Focus
This checkpoint is a specific cell within a larger research grid, designed to investigate the impact of SVD compression on model safety and the efficacy of various component-selection rules for repair. The study aims to quantify how compression alone can degrade safety, specifically by increasing attack-success rates, and to test recovery methods. It is crucial to understand that this model is a research artifact and not a production-ready, general-purpose chat assistant.
Measured Safety Metrics
Key safety metrics measured for this specific checkpoint include:
- AdvBench ASR (HarmBench judge): 0.5900
- StrongREJECT ASR (HarmBench judge): 0.5650
- Macro over-refusal (WildGuard): 0.0278
Intended Use and Limitations
This model is primarily for measuring safety/utility trade-offs under compression. It is important to note that several arms in the research grid, including this one, are deliberately safety-degraded relative to the original Llama-3-8B-Instruct. Users should treat this as an experimental subject and conduct their own evaluations rather than deploying it as a general assistant.