ducthang1703/cbg-llama2-7b-step500
The ducthang1703/cbg-llama2-7b-step500 is a 7 billion parameter language model, derived from Meta's Llama-2-7b-chat-hf, specifically fine-tuned for defense against adversarial attacks using the CBG v2 method. It was trained on a dataset combining RepNoise BeaverTails refusals and Alpaca benign rows over 500 steps. This model is optimized to enhance robustness and safety, making it suitable for applications requiring strong defense mechanisms against harmful or adversarial prompts.
Loading preview...
Overview
ducthang1703/cbg-llama2-7b-step500 is a 7 billion parameter language model, fine-tuned from meta-llama/Llama-2-7b-chat-hf. Its primary purpose is to implement a CBG v2 defense mechanism, enhancing the model's robustness against adversarial inputs. The training involved 500 steps on a specialized dataset that includes both RepNoise BeaverTails refusals (4,000 rows) and Alpaca benign rows (4,000 rows).
Key Capabilities
- Adversarial Defense: Specifically trained with the CBG v2 method to defend against harmful or adversarial prompts.
- Robustness: Aims to improve the model's ability to resist generating undesirable outputs when faced with challenging inputs.
- Llama 2 Base: Benefits from the foundational capabilities of the Llama 2 architecture.
Training Details
The model was trained using specific settings including a learning rate of 2e-05, a cosine scheduler, and a batch size of 8/32. It utilizes bfloat16 for weights dtype and has a maximum length of 256 for both safety and benign contexts. The training process involved specific geometry parameters (alpha, beta, k, radius r, lazy_period, lazy_mode, num_sign_draws, geometry_clip_factor) designed for the CBG v2 defense.
Good For
- Safety-critical applications: Where resistance to adversarial prompting is paramount.
- Research in AI safety: For exploring and evaluating defense mechanisms in large language models.
- Content moderation systems: To enhance the filtering of harmful content by making the model more resilient to manipulation.