wxzhang/dpo-selective-redteaming
The wxzhang/dpo-selective-redteaming model is a 7 billion parameter language model fine-tuned from HuggingFaceH4/zephyr-7b-beta. It was trained using a DPO (Direct Preference Optimization) approach, focusing on improving reward metrics related to chosen and rejected responses. This model is specifically designed for tasks involving selective red teaming, aiming to enhance safety and alignment by optimizing for preferred outputs.
Loading preview...
Model Overview
The wxzhang/dpo-selective-redteaming is a 7 billion parameter language model derived from the HuggingFaceH4/zephyr-7b-beta architecture. This model has undergone fine-tuning using a Direct Preference Optimization (DPO) approach, which aims to align the model's outputs with human preferences by optimizing against chosen and rejected responses.
Key Training Details
- Base Model: HuggingFaceH4/zephyr-7b-beta
- Optimization Method: Direct Preference Optimization (DPO)
- Training Hyperparameters:
- Learning Rate: 5e-07
- Batch Size: 2 (train), 8 (eval)
- Gradient Accumulation Steps: 4
- Optimizer: Adam with betas=(0.9, 0.999)
- Epochs: 1
Performance Metrics (on evaluation set)
During training, the model achieved specific reward metrics, including:
- Rewards/chosen: -0.4246
- Rewards/rejected: -0.4942
- Rewards/accuracies: 0.5239
- Rewards/margins: 0.0696
Intended Use
This model is primarily intended for applications requiring enhanced safety and alignment through selective red teaming. Its DPO fine-tuning suggests a focus on generating more preferred and safer responses, making it suitable for scenarios where output quality and alignment are critical.