promotion/Qwen3-8B-RewardedSoups-baseline

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 6, 2026License:llama3.1Architecture:Transformer Featherless Exclusive Cold

The promotion/Qwen3-8B-RewardedSoups-baseline is an 8 billion parameter language model built on the Qwen3-8B backbone. It utilizes a "Rewarded Soups" technique, linearly interpolating multiple DPO experts, each optimized for specific objectives like instruction following, truthfulness, honesty, and helpfulness. This model is designed to enhance general capabilities by combining specialized fine-tuned policies.

Loading preview...

Model Overview

The promotion/Qwen3-8B-RewardedSoups-baseline is an 8 billion parameter language model based on the Qwen/Qwen3-8B backbone. It implements the "Rewarded Soups" method, which involves training multiple DPO (Direct Preference Optimization) experts, each focused on a distinct objective, and then linearly interpolating their weights. This approach aims to combine the strengths of individual specialized policies into a single, more robust model.

Key Capabilities & Features

  • Multi-objective Optimization: Incorporates DPO experts for instruction following, truthfulness, honesty, and helpfulness.
  • Rewarded Soups Technique: Leverages linear interpolation of expert weights, chosen on a small validation set, to balance different objectives.
  • Backbone: Built upon the Qwen3-8B architecture.
  • Training Details: Trained with 300 optimizer updates and a global batch size of 16.

Evaluation & Use Cases

This model's performance is reported in the Nash Bargaining Preference Optimization (NBPO) paper, specifically in Table 1 for general-capability evaluation. It was assessed using an independent objective-wise win rate against a common reference, judged by Llama-3.3-70B-Instruct on held-out prompts. This model is suitable for applications requiring a balanced performance across multiple ethical and instructional objectives, aiming for improved general utility compared to a single-objective fine-tune.