ducthang1703/cbg-llama2-7b-beta0
The ducthang1703/cbg-llama2-7b-beta0 is a 7 billion parameter language model based on Meta's Llama-2-7b-chat-hf, fine-tuned with a full-parameter CBG v2 defense. It was trained on a specific dataset including RepNoise BeaverTails refusals and Alpaca benign rows to enhance safety and refusal capabilities. This model is designed for applications requiring robust defense against harmful prompts, utilizing a direct "Question: {prompt}\nAnswer:" format.
Loading preview...
Overview
ducthang1703/cbg-llama2-7b-beta0 is a 7 billion parameter language model derived from meta-llama/Llama-2-7b-chat-hf. This model has undergone a full-parameter CBG v2 defense training, specifically designed to improve its safety and refusal mechanisms against problematic inputs. It was trained using a unique data recipe that includes 4,000 defense rows from RepNoise BeaverTails refusals and 4,000 benign rows from Alpaca.
Key Capabilities
- Enhanced Safety: Incorporates a CBG v2 defense mechanism to better handle and refuse harmful or undesirable prompts.
- Specific Prompt Format: Designed to operate with a direct
Question: {prompt}\nAnswer:format, as used during its training. - Llama 2 Foundation: Benefits from the strong base capabilities of the Llama 2 7B chat model.
Training Details
The model was trained for 500 steps with a learning rate of 2e-05 and a cosine scheduler. Key training parameters include a batch size of 8/32, alpha of 1.0, beta of 0.0, and a radius r of 1.1e-07. The maximum length for both safety and benign sequences was set to 256 tokens, and weights were stored in bfloat16 format. Training was conducted on NVIDIA H200 NVL GPUs.
Good for
- Applications requiring a Llama 2-based model with improved safety and refusal capabilities.
- Scenarios where a direct question-answer prompt format is preferred.
- Research into defense mechanisms for large language models.