Jeesup/svd-safety-l2_swift_remove30

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7BQuant:FP8Context Size:4kPublished:Sep 16, 2026License:llama2Architecture:Transformer Open Weights Featherless Exclusive Cold

Jeesup/svd-safety-l2_swift_remove30 is a 7 billion parameter Llama-2-7b-chat checkpoint compressed using Swift-SVD, retaining 70% of its original parameters. This model is a research artifact designed to study how SVD compression impacts safety behavior and to test recovery methods. It is specifically intended for experimental evaluation of safety/utility trade-offs under compression, rather than as a general-purpose chat model.

Loading preview...

Overview

This model, svd-safety-l2_swift_remove30, is a 7 billion parameter Llama-2-7b-chat checkpoint that has undergone significant compression. Developed by Jeesup, it utilizes Swift-SVD with dynamic rank allocation (alpha 0.6, 256 x 2048 WikiText2 calibration) to remove 30% of its dense parameters, resulting in a model that retains 70% of its original parameter count. The compression process was followed by a stage-2 LoRA recovery using SVD-LLM.

Key Characteristics

  • Base Model: meta-llama/Llama-2-7b-chat-hf
  • Compression Method: Swift-SVD, reducing parameters to 69.98% of the original.
  • Recovery: SVD-LLM's stage-2 LoRA (sequential U then V, alpaca-cleaned, r=8, alpha=16, 2 epochs per half, lr 0.0001, batch 64, cutoff 256).
  • Measured Metrics: Achieves an AdvBench ASR of 0.1019 and StrongREJECT ASR of 0.0703 (both via HarmBench judge), with a WikiText-2 perplexity of 9.8411.

Intended Use

This model is a research artifact from a study on how SVD compression affects safety and how best to repair it. It is explicitly not a general-purpose chat model. Its primary purpose is to measure safety/utility trade-offs under compression, with some arms of the study, including this one, being deliberately safety-degraded relative to the original Llama-2-7b-chat. Users should treat it as an experimental subject for evaluation rather than a deployable assistant.