wvnvwn/llama-2-13b-chat-hf-lr5e-5-safeinstr-0.1

TEXT GENERATIONConcurrent Unit Cost:1Model Size:13BQuant:FP8Context Size:4kPublished:Apr 30, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The wvnvwn/llama-2-13b-chat-hf-lr5e-5-safeinstr-0.1 model is a 13 billion parameter language model developed by wvnvwn, based on the Llama-3.2-3B-Instruct architecture. It has been fine-tuned using the Safety Neuron-Tuned (SN-Tune) method on safety alignment data, specifically the Circuit Breakers dataset. This model is optimized for enhanced safety alignment while preserving general capabilities, making it suitable for applications requiring robust content moderation and responsible AI interactions.

Loading preview...

Model Overview

This model, wvnvwn/llama-2-13b-chat-hf-lr5e-5-safeinstr-0.1, is a 13 billion parameter language model derived from the meta-llama/Llama-3.2-3B-Instruct base model. Its primary differentiator is the application of Safety Neuron-Tuned (SN-Tune) fine-tuning, a method developed by wvnvwn to enhance safety alignment.

Key Capabilities & Features

  • Enhanced Safety Alignment: Fine-tuned specifically on the Circuit Breakers dataset using SN-Tune to improve safety responses.
  • Parameter-Efficient Fine-tuning: The SN-Tune method selectively fine-tunes only a small set of "safety neurons," freezing other parameters. This minimizes the computational cost of safety alignment.
  • Preservation of General Capabilities: By focusing only on safety-critical neurons, the model aims to maintain the general language understanding and generation abilities of its base model.
  • Llama Architecture: Built upon the Llama-3.2-3B-Instruct foundation, providing a robust and widely recognized architecture.

What Makes This Model Different?

Unlike general-purpose instruction-tuned models, this model's unique SN-Tune approach specifically targets and modifies only the neurons responsible for safety. This allows for a highly focused and efficient method of mitigating harmful outputs without extensively retraining the entire model. It offers a balance between performance and safety, making it a strong candidate for applications where responsible AI behavior is paramount.

Recommended Use Cases

  • Content Moderation: Ideal for filtering or flagging potentially unsafe content in user-generated text.
  • Responsible AI Applications: Suitable for chatbots, virtual assistants, or any system where preventing harmful or biased responses is critical.
  • Safety Research: Can serve as a baseline for further research into safety alignment techniques and neuron-level modifications.

This model is licensed under the Apache 2.0 License.