Jeesup/svd-safety-l2_basisresmix_remove50

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7BQuant:FP8Context Size:4kPublished:Sep 23, 2026License:llama2Architecture:Transformer Open Weights Featherless Exclusive Cold

Jeesup/svd-safety-l2_basisresmix_remove50 is a Llama-2-7b-chat checkpoint compressed using Basis Sharing, reducing its parameters by 50%. This 7 billion parameter model, with a 4096 token context length, is a research artifact designed to study how SVD compression impacts safety behavior and to test recovery methods. It is not intended as a general-purpose chat model but rather as an experimental subject for analyzing safety/utility trade-offs under compression.

Loading preview...

Model Overview

This model, svd-safety-l2_basisresmix_remove50, is a research artifact derived from meta-llama/Llama-2-7b-chat-hf. It has been significantly compressed using Basis Sharing, an ICLR 2025 technique that shares bases over groups of two adjacent layers, resulting in a 50% reduction in dense parameters. The model was then recovered through a coefficient-only LoRA fine-tune, maintaining the parameter budget and shared bases.

Key Characteristics

  • Base Model: Llama-2-7b-chat-hf
  • Compression Method: Basis Sharing, reducing parameters by 50.0%
  • Recovery Method: LoRA (r=8) on per-layer coefficients, 2 epochs, lr 0.0001, batch 64, alpaca-cleaned dataset.
  • Measured Metrics:
    • AdvBench ASR (HarmBench judge): 0.1808
    • StrongREJECT ASR (HarmBench judge): 0.1565
    • Macro over-refusal (WildGuard): 0.1646
    • WikiText-2 perplexity: 13.2626

Intended Use and Limitations

This model is not a general-purpose chat model. Its primary purpose is to serve as an experimental subject within a research study investigating the impact of SVD compression on safety behavior and the effectiveness of various recovery techniques. The model is part of a grid of experiments, and some configurations are deliberately safety-degraded relative to the original Llama-2-7b-chat. Users should treat this checkpoint as a research artifact for measuring safety/utility trade-offs under compression, rather than a deployable assistant.