promotion/Qwen3-8B-MaxMinRLHF-baseline

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 6, 2026License:llama3.1Architecture:Transformer Featherless Exclusive Cold

The promotion/Qwen3-8B-MaxMinRLHF-baseline is an 8 billion parameter language model based on the Qwen3-8B architecture, developed by the creators of MaxMin-RLHF. It is fine-tuned using the MaxMin-RLHF algorithm to optimize for multiple objectives including instruction following, truthfulness, honesty, and helpfulness. This model is designed for general-capability applications where balanced performance across these four objectives is critical, offering a robust baseline for multi-objective optimization.

Loading preview...

Model Overview

The promotion/Qwen3-8B-MaxMinRLHF-baseline is an 8 billion parameter language model built upon the Qwen3-8B backbone. It incorporates the MaxMin-RLHF (Chakraborty et al., ICML 2024) Algorithm 1, which iteratively optimizes for a panel of four distinct objectives: instruction following, truthfulness, honesty, and helpfulness. The training involved three alternating rounds of 100 updates, focusing on the objective with the lowest reference-standardized utility, totaling 300 optimizer updates.

Key Capabilities

  • Multi-objective Optimization: Specifically trained to balance performance across instruction following, truthfulness, honesty, and helpfulness.
  • Robust Baseline: Serves as a general-capability model, evaluated for its balanced performance across the specified objectives.
  • RLHF Integration: Utilizes the MaxMin-RLHF algorithm for preference optimization, aiming for equitable performance across different utility metrics.

Good For

  • General-purpose applications: Where a balanced and robust performance across multiple ethical and functional objectives is desired.
  • Research in RLHF: As a reference implementation for the MaxMin-RLHF algorithm.
  • Applications requiring balanced outputs: Particularly in scenarios where instruction adherence, factual accuracy, ethical considerations, and user assistance are all important.