samarthraina/Llama-3-8B-D-STEER-OpenHermes-IT

TEXT GENERATIONPricing:Input $0.37 / Cached $0.074 / Output $0.38Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Aug 19, 2026License:llama3Architecture:Transformer Featherless Exclusive Cold

The samarthraina/Llama-3-8B-D-STEER-OpenHermes-IT is an 8 billion parameter instruction-tuned causal language model based on Meta Llama 3 8B, developed by samarthraina. Fine-tuned on the OpenHermes-2.5 dataset, this model serves as the unsteered reference checkpoint for D-STEER research into activation steering for AI safety. It is specifically intended for AI safety research, particularly for measuring and comparing alignment behavior and as a starting point for activation-steering experiments.

Loading preview...

D-STEER Llama-3-8B OpenHermes — IT checkpoint

This model is an 8 billion parameter instruction-tuned checkpoint derived from Meta Llama 3 8B, fine-tuned using the OpenHermes-2.5 dataset. It is released as part of the D-STEER research program, which investigates activation steering to control safety behaviors in language models.

Key Characteristics

  • Base Model: Meta Llama 3 8B.
  • Instruction-Tuning: Supervised fine-tuning on the OpenHermes-2.5 dataset.
  • Format: Self-contained merged causal language model, requiring no PEFT adapter.
  • Precision: Stored in float16.

Intended Use

This model is primarily intended for AI safety research. It functions as the instruction-tuned starting point (IT checkpoint) in D-STEER experiments. Researchers use it as a reference against which harmlessness is installed via activation steering, and as the unsteered endpoint for measuring and comparing alignment behavior.

Limitations and Safety

It is crucial to note that this IT checkpoint has no additional harmlessness training and is designed to be the less safe member of its model pair. It may comply with harmful requests more often than its preference-tuned counterpart. It is not evaluated or hardened for production deployment and inherits the limitations, biases, and knowledge cutoff of Meta Llama 3 8B and the OpenHermes-2.5 data. Its behavior in languages other than English is untested.

Matched Checkpoint

Its preference-tuned counterpart, used in D-STEER for installing safety behavior, is available as samarthraina/Llama-3-8B-D-STEER-OpenHermes-DPO.