kmseong/qwen2_5_32b_instruct-lr5e-5-gsm8k-resta-gamma0.3
The kmseong/WaRP-Safety-Llama3_8B_Instruct is an 8 billion parameter Llama 3.1 Instruct model fine-tuned for safety alignment using the Weight space Rotation Process (WaRP). This model focuses on maintaining refusal capabilities for harmful requests while improving utility on reasoning tasks like GSM8K. It is designed to balance safety and performance, making it suitable for applications requiring robust safety mechanisms.
Loading preview...
Overview
This model, kmseong/WaRP-Safety-Llama3_8B_Instruct, is an 8 billion parameter instruction-tuned Llama 3.1 model developed by Min-Seong Kim. It has been fine-tuned using a novel Safety-First Weight space Rotation Process (WaRP), a three-phase pipeline designed to enhance safety alignment while preserving utility.
Key Capabilities
- Enhanced Safety Alignment: Utilizes a sophisticated WaRP method to protect safety mechanisms through gradient masking during fine-tuning.
- Refusal Capability: Maintains the ability to refuse harmful requests effectively.
- Improved Utility: Demonstrates improved performance on reasoning tasks, specifically fine-tuned using the GSM8K dataset.
- Balanced Safety-Utility Tradeoff: Aims to achieve a strong balance between safety and general task performance.
- Training Methodology: Involves Basis Construction (identifying important neurons), Importance Scoring (calculating gradient-based scores), and Incremental Learning (fine-tuning with gradient masking).
Good For
- Applications requiring a strong emphasis on safety and ethical AI behavior.
- Use cases where refusal of harmful content is critical.
- Tasks that benefit from a model with improved reasoning capabilities while maintaining safety.
- Developers looking for a Llama 3.1 Instruct variant with explicit safety alignment through a specialized training process.