promotion/Llama-3.1-8B-SurplusMaxmin-baseline
Llama-3.1-8B-SurplusMaxmin-baseline is an 8 billion parameter language model based on Meta's Llama-3.1-8B-Instruct, fine-tuned using the SurplusMaxMin method on the UltraFeedback panel. This model focuses on optimizing instruction following, truthfulness, honesty, and helpfulness. It was developed as part of research into Nash Bargaining Preference Optimization (NBPO) to explore objective weighting strategies. The model is designed for applications requiring balanced performance across multiple ethical and instructional criteria.
Loading preview...
Model Overview
promotion/Llama-3.1-8B-SurplusMaxmin-baseline is an 8 billion parameter language model derived from meta-llama/Llama-3.1-8B-Instruct. It was fine-tuned using the SurplusMaxMin method, a technique for objective weighting within preference optimization frameworks. This model is specifically designed to balance performance across several key objectives: instruction following, truthfulness, honesty, and helpfulness.
Key Characteristics
- Base Model:
meta-llama/Llama-3.1-8B-Instruct - Training Method: SurplusMaxMin, a game-aggregation control approach for objective weighting.
- Feedback Panel: Utilizes the UltraFeedback dataset for training.
- Optimized Objectives: Focuses on improving instruction following, truthfulness, honesty, and helpfulness simultaneously.
- Training Budget: Trained with 300 optimizer updates and a global batch size of 16.
- Evaluation: Performance is assessed by independent objective-wise win rates against the reference policy, judged by
Llama-3.3-70B-Instructon held-out prompts.
Use Cases
This model is suitable for applications where a balanced emphasis on multiple ethical and instructional criteria is important. It is particularly relevant for research into preference optimization and multi-objective alignment, offering a baseline for comparing different objective weighting strategies.