Jeesup/svd-safety-l2_swift_remove20

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7BQuant:FP8Context Size:4kPublished:Sep 16, 2026License:llama2Architecture:Transformer Open Weights Featherless Exclusive Cold

Jeesup/svd-safety-l2_swift_remove20 is a 7 billion parameter Llama-2-7b-chat checkpoint compressed using Swift-SVD to 80% of its original parameters, then recovered with SVD-LLM's stage-2 LoRA. This model is a research artifact designed to study how SVD compression impacts safety behavior and the effectiveness of recovery methods. It is specifically intended for measuring safety/utility trade-offs under compression, rather than serving as a general-purpose chat model.

Loading preview...

Overview

Jeesup/svd-safety-l2_swift_remove20 is a research-oriented model derived from meta-llama/Llama-2-7b-chat-hf. It underwent Swift-SVD compression, reducing its parameters by 20% (resulting in 80% of dense parameters), followed by a stage-2 LoRA recovery process. The compression used dynamic rank allocation with an alpha of 0.6 and WikiText2 calibration.

Key Characteristics

  • Base Model: meta-llama/Llama-2-7b-chat-hf
  • Compression Method: Swift-SVD (dynamic rank allocation, alpha 0.6, 256 x 2048 WikiText2 calibration)
  • Parameter Reduction: 20% of parameters removed, resulting in 0.7999 parameter fraction.
  • Recovery Method: SVD-LLM's stage-2 LoRA (sequential U then V, alpaca-cleaned, r=8, alpha=16, 2 epochs per half, lr 0.0001, batch 64, cutoff 256).

Measured Performance (Research Context)

  • AdvBench ASR (HarmBench judge): 0.0135
  • StrongREJECT ASR (HarmBench judge): 0.0319
  • Macro over-refusal (WildGuard): 0.3632
  • WikiText-2 perplexity: 8.7380

Intended Use and Limitations

This model is a research artifact from a study on how SVD compression affects safety and how different component-selection rules repair it. It is not a general-purpose chat model and is deliberately safety-degraded relative to the original Llama-2-7b-chat due to compression. Users should treat it as an experimental subject for evaluating safety/utility trade-offs under compression, rather than a deployable assistant.