kmseong/llama-2-7b-chat-hf-arc-rsn-tuned-lr5e-5

TEXT GENERATIONConcurrent Unit Cost:1Model Size:7BQuant:FP8Context Size:4kPublished:May 3, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The kmseong/llama-2-7b-chat-hf-arc-rsn-tuned-lr5e-5 is a 7 billion parameter Llama-2-chat-hf model, fine-tuned by kmseong using a Safety Neuron-Tuning (SN-Tune) approach. This method selectively fine-tunes only safety-critical neurons on safety alignment data, enhancing safety without significantly impacting general capabilities. It is designed to provide improved safety alignment compared to its base model, making it suitable for applications requiring robust safety features.

Loading preview...

Model Overview

This model, kmseong/llama-2-7b-chat-hf-arc-rsn-tuned-lr5e-5, is a 7 billion parameter variant of the Llama-2-chat-hf architecture. It has been specifically fine-tuned by kmseong using a novel Safety Neuron-Tuning (SN-Tune) method. The base model for this fine-tuning is meta-llama/Llama-3.2-3B-Instruct.

Key Features & SN-Tune Methodology

The core differentiator of this model is its SN-Tune approach, which involves:

  • Selective Fine-tuning: Identifies and targets a small subset of "safety neurons" crucial for alignment.
  • Parameter Efficiency: Freezes all non-safety parameters, fine-tuning only the identified safety neurons.
  • Enhanced Safety: Trained on the "Circuit Breakers" dataset, focusing on safety alignment data.

This methodology aims to achieve improved safety alignment while preserving the general capabilities of the base model and ensuring parameter-efficient fine-tuning.

Use Cases

This model is particularly well-suited for applications where:

  • Safety is paramount: Its SN-Tune process is designed to enhance safety alignment.
  • Resource efficiency is desired: The selective fine-tuning minimizes computational overhead compared to full model fine-tuning.
  • General conversational abilities are needed: It retains the conversational capabilities of the Llama-2-chat-hf base model, with added safety enhancements.