Jeesup/svd-safety-mis7_swift_remove20
Jeesup/svd-safety-mis7_swift_remove20 is a 7 billion parameter Mistral-7B-Instruct-v0.2 checkpoint, compressed using Swift-SVD to 80% of its original parameters and then recovered with SVD-LLM's stage-2 LoRA. This model serves as a research artifact to study how SVD compression impacts safety behavior and the effectiveness of different component-selection rules for repair. It is specifically designed for experimental evaluation of safety/utility trade-offs under compression, rather than as a general-purpose chat model.
Loading preview...
Overview
Jeesup/svd-safety-mis7_swift_remove20 is a 7 billion parameter model derived from mistralai/Mistral-7B-Instruct-v0.2. It has undergone significant compression using Swift-SVD, reducing its parameters to 80.0% of the original, followed by recovery using SVD-LLM's stage-2 LoRA. This process involved dynamic rank allocation with an alpha of 0.6 and WikiText2 calibration.
Key Characteristics
- Base Model: Mistral-7B-Instruct-v0.2
- Compression Method: Swift-SVD (20.00% of parameters removed)
- Recovery Method: SVD-LLM's stage-2 LoRA (alpaca-cleaned, r=8, alpha=16)
- Context Length: 4096 tokens
- Measured Metrics:
- AdvBench ASR (HarmBench judge): 0.5692
- StrongREJECT ASR (HarmBench judge): 0.4185
- Macro over-refusal (WildGuard): 0.1051
- WikiText-2 perplexity: 7.4204
Intended Use and Limitations
This model is a research artifact specifically created to measure safety/utility trade-offs under compression. It is part of a study quantifying how compression degrades safety and testing recovery methods. Users should be aware that this model is deliberately safety-degraded relative to the original Mistral-7B-Instruct-v0.2. It is not intended as a deployable general-purpose chat model but rather as an experimental subject for research purposes. Users are advised to evaluate it thoroughly before drawing conclusions.