kmseong/qwen2_5_32b-instruct-warp-lr5e-5
The kmseong/qwen2_5_32b-instruct-warp-lr5e-5 model is a 32.8 billion parameter instruction-tuned language model based on the Qwen2.5 architecture. It is fine-tuned using a Weight space Rotation Process (WaRP) for safety alignment, focusing on maintaining refusal capabilities for harmful requests while improving utility on reasoning tasks. This model is designed to offer a balanced safety-utility tradeoff, making it suitable for applications requiring robust safety features alongside general language understanding.
Loading preview...
Overview
This model, kmseong/qwen2_5_32b-instruct-warp-lr5e-5, is a 32.8 billion parameter instruction-tuned variant of the Qwen2.5 architecture. It has been specifically fine-tuned using a novel Weight space Rotation Process (WaRP) to enhance safety alignment. The WaRP method involves a 3-phase pipeline: Basis Construction, Importance Scoring, and Incremental Learning, which collectively aim to protect safety mechanisms while improving utility.
Key Capabilities
- Safety Alignment: Utilizes the WaRP method to maintain refusal capabilities for harmful requests through gradient masking.
- Utility Improvement: Designed to improve performance on reasoning tasks, such as GSM8K, while preserving safety.
- Balanced Performance: Achieves a balance between safety and utility, making it suitable for sensitive applications.
- Instruction Following: Inherits strong instruction-following capabilities from its Qwen2.5-Instruct base.
Good For
- Applications requiring robust safety: Ideal for use cases where preventing harmful or inappropriate outputs is critical.
- General-purpose instruction following: Can be used for a wide range of NLP tasks where a balance of safety and performance is desired.
- Research into safety alignment techniques: Provides a practical example of the WaRP method's application in LLM safety.
Limitations
As with all fine-tuned models, users should evaluate outputs for their specific use cases and consider additional safety measures. The model's license follows that of the base Qwen2.5 model.