kmseong/qwen2_5_32b-instruct-gsm8k-rsn-tuned-lr5e-5
This model is a 32.8 billion parameter instruction-tuned variant of the Llama-3.2-3B-Instruct architecture, developed by kmseong. It has been fine-tuned using the Safety Neuron Tuning (SN-Tune) method on safety alignment data, specifically targeting enhanced safety without significantly impacting general capabilities. The model is optimized for improved safety alignment through parameter-efficient fine-tuning of critical safety neurons. It supports a context length of 32768 tokens.
Loading preview...
Model Overview
This model, kmseong/qwen2_5_32b-instruct-gsm8k-rsn-tuned-lr5e-5, is a specialized version of the Llama-3.2-3B-Instruct base model, developed by kmseong. It features 32.8 billion parameters and a 32768-token context length. The primary differentiator of this model is its fine-tuning using the Safety Neuron Tuning (SN-Tune) method.
Safety Neuron Tuning (SN-Tune)
SN-Tune is a selective fine-tuning approach designed to enhance safety alignment efficiently. It operates by:
- Detecting a small, critical set of "safety neurons" within the model.
- Freezing all other parameters to preserve general capabilities.
- Fine-tuning only these identified safety neurons on dedicated safety alignment data, such as the Circuit Breakers dataset.
This method aims to provide enhanced safety alignment with minimal impact on the model's general performance and offers a parameter-efficient way to achieve safety improvements.
Key Capabilities
- Enhanced Safety Alignment: Specifically tuned to improve safety responses compared to its base model.
- Parameter-Efficient Fine-tuning: Achieves safety improvements by selectively tuning a small subset of neurons.
- Instruction Following: Inherits instruction-following capabilities from its Llama-3.2-3B-Instruct base.
Use Cases
This model is particularly suitable for applications where:
- Safety and responsible AI are paramount: Ideal for deployments requiring robust safety guardrails.
- Maintaining general capabilities is important: The SN-Tune method aims to prevent degradation of non-safety-related performance.
- Efficient safety updates are desired: The selective fine-tuning approach allows for targeted safety enhancements.