promotion/Llama-3.1-8B-SafeRLHF-AbsoluteMaxMin-baseline

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 5, 2026License:llama3.1Architecture:Transformer Featherless Exclusive Cold

Llama-3.1-8B-SafeRLHF-AbsoluteMaxMin-baseline is an 8 billion parameter language model developed by promotion, based on Meta's Llama-3.1-8B-Instruct. This model is fine-tuned using the AbsoluteMaxMin objective within the SafeRLHF framework, focusing on balancing helpfulness and harmlessness. It is designed for applications requiring robust safety alignment and balanced performance across multiple objectives.

Loading preview...

Overview

This model, Llama-3.1-8B-SafeRLHF-AbsoluteMaxMin-baseline, is an 8 billion parameter language model derived from meta-llama/Llama-3.1-8B-Instruct. It was developed by promotion as part of the SafeRLHF panel, specifically utilizing the AbsoluteMaxMin objective for alignment.

Key Capabilities

  • Safety Alignment: Fine-tuned with a strong emphasis on both helpfulness and harmlessness objectives.
  • Objective Balancing: Employs the AbsoluteMaxMin method to achieve a balanced performance across multiple objectives, as detailed in the Nash Bargaining Preference Optimization (NBPO) research.
  • Controlled Training: Trained with a matched game-aggregation control, ensuring the same response pool, temperatures, disagreement point, and optimizer budget as other methods like NBPO.

Training and Evaluation

  • Training Budget: Underwent 300 optimizer updates with a global batch size of 16.
  • Evaluation: Performance is assessed by an independent objective-wise win rate against a common reference, judged by Llama-3.3-70B-Instruct on prompt-disjoint held-out prompts.

Good For

  • Applications where balancing helpfulness and harmlessness is critical.
  • Research and development in safety-aligned reinforcement learning from human feedback (RLHF).
  • Use cases requiring a robustly aligned 8B parameter model.