samarthraina/Llama-3-8B-D-STEER-OpenHermes-DPO
The samarthraina/Llama-3-8B-D-STEER-OpenHermes-DPO is an 8 billion parameter Llama-3 based model, fine-tuned using Direct Preference Optimization (DPO) on the Anthropic HH-RLHF dataset for harmlessness. This model is specifically designed for AI safety research, serving as a preference-tuned endpoint in a matched pair to estimate and evaluate harmlessness directions in activation space. It is built upon the Meta Llama 3 8B foundation model and instruction-tuned with OpenHermes-2.5 data, offering a 8192 token context length.
Loading preview...
Model Overview
This model, samarthraina/Llama-3-8B-D-STEER-OpenHermes-DPO, is an 8 billion parameter Llama-3 based checkpoint developed by samarthraina. It has undergone Direct Preference Optimization (DPO) specifically targeting harmlessness, utilizing the Anthropic HH-RLHF dataset. The model is built on the meta-llama/Meta-Llama-3-8B foundation and instruction-tuned with teknium/OpenHermes-2.5 data, featuring a context length of 8192 tokens.
Key Characteristics
- Harmlessness-Oriented DPO: Fine-tuned using Direct Preference Optimization on the Anthropic HH-RLHF dataset to enhance harmlessness behavior.
- Matched Pair for Research: This DPO checkpoint is part of a matched pair with
samarthraina/Llama-3-8B-D-STEER-OpenHermes-IT, allowing researchers to estimate and evaluate harmlessness directions in activation space. - Llama-3 Architecture: Inherits the robust Llama-3 8B architecture, providing a strong base for its specialized fine-tuning.
- OpenHermes-2.5 Instruction Tuning: Benefits from instruction tuning on the OpenHermes-2.5 dataset, contributing to its general conversational abilities before the preference optimization stage.
Intended Use Cases
- AI Safety Research: Primarily released for research into AI safety, particularly for studying and quantifying harmlessness in language models through activation steering.
- Activation Steering Studies: Ideal for experiments involving the estimation and manipulation of safety behaviors in activation space.
Limitations
- Harmlessness training reduces, but does not eliminate, harmful completions.
- The preference stage focuses solely on harmlessness and does not guarantee general capability, factuality, or robustness improvements.
- May exhibit increased refusal rates on benign requests due to its harmlessness optimization.
- Not evaluated for production deployment and inherits limitations from its base models and training data.