samarthraina/Llama-3-8B-D-STEER-OpenHermes-DPO

TEXT GENERATIONPricing:Input $0.37 / Cached $0.074 / Output $0.38Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Aug 19, 2026License:llama3Architecture:Transformer Featherless Exclusive Cold

The samarthraina/Llama-3-8B-D-STEER-OpenHermes-DPO is an 8 billion parameter Llama-3 based model, fine-tuned using Direct Preference Optimization (DPO) on the Anthropic HH-RLHF dataset for harmlessness. This model is specifically designed for AI safety research, serving as a preference-tuned endpoint in a matched pair to estimate and evaluate harmlessness directions in activation space. It is built upon the Meta Llama 3 8B foundation model and instruction-tuned with OpenHermes-2.5 data, offering a 8192 token context length.

Loading preview...

Model Overview

This model, samarthraina/Llama-3-8B-D-STEER-OpenHermes-DPO, is an 8 billion parameter Llama-3 based checkpoint developed by samarthraina. It has undergone Direct Preference Optimization (DPO) specifically targeting harmlessness, utilizing the Anthropic HH-RLHF dataset. The model is built on the meta-llama/Meta-Llama-3-8B foundation and instruction-tuned with teknium/OpenHermes-2.5 data, featuring a context length of 8192 tokens.

Key Characteristics

  • Harmlessness-Oriented DPO: Fine-tuned using Direct Preference Optimization on the Anthropic HH-RLHF dataset to enhance harmlessness behavior.
  • Matched Pair for Research: This DPO checkpoint is part of a matched pair with samarthraina/Llama-3-8B-D-STEER-OpenHermes-IT, allowing researchers to estimate and evaluate harmlessness directions in activation space.
  • Llama-3 Architecture: Inherits the robust Llama-3 8B architecture, providing a strong base for its specialized fine-tuning.
  • OpenHermes-2.5 Instruction Tuning: Benefits from instruction tuning on the OpenHermes-2.5 dataset, contributing to its general conversational abilities before the preference optimization stage.

Intended Use Cases

  • AI Safety Research: Primarily released for research into AI safety, particularly for studying and quantifying harmlessness in language models through activation steering.
  • Activation Steering Studies: Ideal for experiments involving the estimation and manipulation of safety behaviors in activation space.

Limitations

  • Harmlessness training reduces, but does not eliminate, harmful completions.
  • The preference stage focuses solely on harmlessness and does not guarantee general capability, factuality, or robustness improvements.
  • May exhibit increased refusal rates on benign requests due to its harmlessness optimization.
  • Not evaluated for production deployment and inherits limitations from its base models and training data.