kmseong/llama2_7b-base-CB_SSFT-lr3e-5

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7BQuant:FP8Context Size:4kPublished:Sep 14, 2026License:llama3.1Architecture:Transformer Featherless Exclusive Cold

The kmseong/llama2_7b-base-CB_SSFT-lr3e-5 model, developed by Min-Seong Kim, is a 7 billion parameter Llama 3.1 8B Instruct model fine-tuned for safety alignment using the Weight space Rotation Process (WaRP). This model focuses on balancing safety and utility, maintaining refusal capabilities for harmful requests while improving performance on reasoning tasks. It is optimized for applications requiring robust safety features alongside general language understanding and generation.

Loading preview...

Overview

The kmseong/WaRP-Safety-Llama3_8B_Instruct model is a Llama 3.1 8B Instruct variant, fine-tuned by Min-Seong Kim using a novel Weight space Rotation Process (WaRP) for enhanced safety alignment. This 7 billion parameter model is designed to improve safety mechanisms while preserving and enhancing utility on reasoning tasks.

Key Capabilities

  • Safety Alignment: Utilizes a 3-phase WaRP pipeline to construct a basis from safety data, score neuron importance, and incrementally learn with gradient masking.
  • Harmful Request Refusal: Specifically trained to maintain refusal capabilities for unsafe or harmful prompts.
  • Utility Preservation: Improves performance on utility tasks, such as mathematical reasoning (demonstrated with GSM8K), by protecting important directions during fine-tuning.
  • Balanced Trade-off: Achieves a balance between safety and utility, ensuring the model remains helpful while being robust against harmful content generation.

Training Details

The model was trained using a three-phase process:

  1. Basis Construction: Identified important neurons in FFN layers using safety data (LibrAI/do-not-answer) and SVD.
  2. Importance Scoring: Calculated gradient-based importance scores to generate masks for critical directions.
  3. Incremental Learning: Fine-tuned on utility data (openai/gsm8k) with gradient masking to protect safety-critical directions, enhancing utility without compromising safety.

Good For

This model is suitable for applications where a strong emphasis on safety and responsible AI is paramount, particularly in scenarios requiring:

  • Content moderation and filtering.
  • Applications needing robust refusal of harmful or inappropriate requests.
  • General language generation and understanding where a balance between helpfulness and safety is crucial.