promotion/Qwen3-8B-MaxMinRLHF-baseline
The promotion/Qwen3-8B-MaxMinRLHF-baseline is an 8 billion parameter language model based on the Qwen3-8B architecture, developed by the creators of MaxMin-RLHF. It is fine-tuned using the MaxMin-RLHF algorithm to optimize for multiple objectives including instruction following, truthfulness, honesty, and helpfulness. This model is designed for general-capability applications where balanced performance across these four objectives is critical, offering a robust baseline for multi-objective optimization.
Loading preview...
Model Overview
The promotion/Qwen3-8B-MaxMinRLHF-baseline is an 8 billion parameter language model built upon the Qwen3-8B backbone. It incorporates the MaxMin-RLHF (Chakraborty et al., ICML 2024) Algorithm 1, which iteratively optimizes for a panel of four distinct objectives: instruction following, truthfulness, honesty, and helpfulness. The training involved three alternating rounds of 100 updates, focusing on the objective with the lowest reference-standardized utility, totaling 300 optimizer updates.
Key Capabilities
- Multi-objective Optimization: Specifically trained to balance performance across instruction following, truthfulness, honesty, and helpfulness.
- Robust Baseline: Serves as a general-capability model, evaluated for its balanced performance across the specified objectives.
- RLHF Integration: Utilizes the MaxMin-RLHF algorithm for preference optimization, aiming for equitable performance across different utility metrics.
Good For
- General-purpose applications: Where a balanced and robust performance across multiple ethical and functional objectives is desired.
- Research in RLHF: As a reference implementation for the MaxMin-RLHF algorithm.
- Applications requiring balanced outputs: Particularly in scenarios where instruction adherence, factual accuracy, ethical considerations, and user assistance are all important.