promotion/Llama-3.1-8B-SafeRLHF-Panacea-baseline

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 5, 2026License:llama3.1Architecture:Transformer Featherless Exclusive Cold

Llama-3.1-8B-SafeRLHF-Panacea-baseline is an 8 billion parameter model based on the Llama-3.1-8B-Instruct backbone, fine-tuned using the Panacea method with DPO and linear scalarization. It is specifically optimized for SafeRLHF objectives, focusing on improving both helpfulness and harmlessness. This model is designed for applications requiring robust performance in safety-critical conversational AI.

Loading preview...

Model Overview

Llama-3.1-8B-SafeRLHF-Panacea-baseline is an 8 billion parameter language model developed by promotion, built upon the meta-llama/Llama-3.1-8B-Instruct backbone. It incorporates the Panacea method (Zhong et al., NeurIPS 2024) with a DPO (Direct Preference Optimization) procedure and linear scalarization. The fine-tuning process utilized SVD-LoRA with k=8 preference-agnostic singular values, injecting the preference vector as additional singular values.

Key Characteristics

  • SafeRLHF Objectives: Primarily optimized for helpfulness and harmlessness, making it suitable for applications where safety and beneficial interactions are paramount.
  • Training Details: Underwent 300 optimizer updates with a global batch size of 16.
  • Evaluation: Performance is assessed using an independent objective-wise win rate against a common reference, judged by Llama-3.3-70B-Instruct on prompt-disjoint held-out prompts.

Intended Use Cases

This model is particularly well-suited for:

  • Safety-critical AI applications: Where mitigating harmful outputs and ensuring helpful responses are crucial.
  • Conversational AI: For chatbots and virtual assistants that require a strong emphasis on ethical and beneficial interactions.
  • Research in RLHF: As a baseline or comparison point for further studies in safe reinforcement learning from human feedback.