Jeesup/svd-safety-l3_remove50_swapgapiter_b010_r02
Jeesup/svd-safety-l3_remove50_swapgapiter_b010_r02 is an 8 billion parameter Llama-3-8B-Instruct checkpoint, compressed to 50% of its original parameters using SVD-LLM. This model is a research artifact designed to study how SVD compression impacts safety behavior and the effectiveness of iterative parameter swapping for repair. It is specifically configured with 2 of 10 rounds of iterative parameter-neutral swap using the 'gap_iter' rule, making it an experimental subject rather than a general-purpose chat model.
Loading preview...
Model Overview
This model, svd-safety-l3_remove50_swapgapiter_b010_r02, is an experimental Llama-3-8B-Instruct checkpoint that has undergone significant compression and subsequent iterative repair. It is a research artifact from a study investigating the effects of SVD-LLM compression on safety behavior and methods to mitigate degradation.
Key Characteristics
- Base Model:
meta-llama/Meta-Llama-3-8B-Instruct. - Compression: Compressed using SVD-LLM, resulting in 50.03% of original parameters removed.
- Repair Mechanism: Features 2 of 10 planned rounds of iterative parameter-neutral swapping, guided by the
gap_iterselection rule, with a per-round budget of 0.1% of dense parameters. - Parameter Count: The resulting model retains approximately 49.97% of the original dense parameters.
- Experimental Nature: This checkpoint is a specific cell within a larger research grid, designed to measure safety/utility trade-offs under compression. It is not intended as a deployable general-purpose chat model.
Measured Safety Metrics
- AdvBench ASR (HarmBench judge): 0.1400
- StrongREJECT ASR (HarmBench judge): 0.1800
- Macro over-refusal (WildGuard): 0.4146
Intended Use and Limitations
This model's primary purpose is to serve as an experimental subject for quantifying safety degradation due to compression and evaluating recovery strategies. It is explicitly noted that several arms in the study, including this one, are deliberately safety-degraded relative to the original Llama-3-8B-Instruct. Users should treat it as a research tool for analysis rather than a production-ready assistant.