promotion/Llama-3.1-8B-AbsoluteMaxmin-baseline
Llama-3.1-8B-AbsoluteMaxmin-baseline is an 8 billion parameter model based on the meta-llama/Llama-3.1-8B-Instruct backbone, fine-tuned using the AbsoluteMaxMin method on the UltraFeedback panel. This model focuses on optimizing instruction following, truthfulness, honesty, and helpfulness. It utilizes a matched game-aggregation control for objective weighting, distinguishing it from other preference optimization techniques. The model is designed for applications requiring balanced performance across multiple ethical and instructional objectives.
Loading preview...
AbsoluteMaxMin on UltraFeedback
This model, promotion/Llama-3.1-8B-AbsoluteMaxmin-baseline, is an 8 billion parameter language model built upon the meta-llama/Llama-3.1-8B-Instruct backbone. It has been fine-tuned using the AbsoluteMaxMin method, a preference optimization technique that differs from others like NBPO primarily in its objective weighting strategy. The training utilized the UltraFeedback panel, focusing on a balanced optimization across four key objectives: instruction following, truthfulness, honesty, and helpfulness.
Key Characteristics
- Backbone:
meta-llama/Llama-3.1-8B-Instruct - Optimization Method: AbsoluteMaxMin, using a matched game-aggregation control for objective weighting.
- Training Data: UltraFeedback panel.
- Objectives: Optimized for instruction following, truthfulness, honesty, and helpfulness.
- Training Budget: 300 optimizer updates with a global batch size of 16.
- Context Length: Supports a context length of 32768 tokens.
Evaluation
The model's performance is evaluated by its independent objective-wise win rate against the common reference model. This assessment is conducted by Llama-3.3-70B-Instruct on prompt-disjoint held-out prompts, considering both presentation orders.
Use Cases
This model is particularly suitable for applications where a balanced and robust performance across multiple ethical and instructional criteria is critical, such as chatbots, content generation, and agents requiring high fidelity in instruction adherence and factual accuracy.