promotion/Llama-3.1-8B-SafeRLHF-RewardedSoups-baseline

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 5, 2026License:llama3.1Architecture:Transformer Featherless Exclusive Cold

The promotion/Llama-3.1-8B-SafeRLHF-RewardedSoups-baseline is an 8 billion parameter language model based on the Llama-3.1-8B-Instruct backbone. Developed by promotion, this model utilizes Rewarded Soups for fine-tuning, combining multiple expert policies for helpfulness and harmlessness. It is optimized for safety and alignment, making it suitable for applications requiring robust and ethically sound AI responses.

Loading preview...

Model Overview

The promotion/Llama-3.1-8B-SafeRLHF-RewardedSoups-baseline is an 8 billion parameter language model built upon the meta-llama/Llama-3.1-8B-Instruct backbone. This model incorporates the Rewarded Soups technique (Rame et al., NeurIPS 2023), which involves linearly interpolating multiple expert policies, each optimized for a specific objective, from a shared initialisation. The interpolation weights are determined using validation prompts.

Key Characteristics

  • Fine-tuning Method: Employs Rewarded Soups within the SafeRLHF framework.
  • Objectives: Optimized for both helpfulness and harmlessness, aiming to produce aligned and safe outputs.
  • Training Details: Underwent 300 optimizer updates with a global batch size of 16.
  • Evaluation: Performance is assessed by independent objective-wise win rates against a common reference, judged by Llama-3.3-70B-Instruct on prompt-disjoint held-out prompts.

Use Cases

This model is particularly well-suited for applications where safety and alignment are paramount. Its fine-tuning for helpfulness and harmlessness makes it a strong candidate for:

  • Content moderation and filtering
  • Customer support and conversational AI requiring ethical responses
  • Educational tools where safe and informative interactions are crucial
  • General-purpose assistants that prioritize responsible AI behavior.