promotion/FixedRefNash-Qwen2.5-7B-WildChecklists-2ep

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 16, 2026Architecture:Transformer Featherless Exclusive Cold

FixedRefNash-Qwen2.5-7B-WildChecklists-2ep is a 7.6 billion parameter research checkpoint model fine-tuned from Qwen/Qwen2.5-7B-Instruct. It utilizes Nash aggregation with a frozen reference policy and per-prompt multipliers, trained on the WildChecklists dataset for specific prompt-based evaluation. This model is part of a comparative study focusing on aggregation rules, offering insights into fixed-reference Nash performance in controlled evaluation settings.

Loading preview...

Overview

This model, FixedRefNash-Qwen2.5-7B-WildChecklists-2ep, is a 7.6 billion parameter research checkpoint derived from Qwen/Qwen2.5-7B-Instruct. It implements a fixed-reference Nash aggregation strategy with per-prompt multipliers, distinguishing it from other models in a comparative study where the primary variable is the aggregation rule. The model was fine-tuned using the "Back to Blackwell" (PROSPER) protocol on the WildChecklists dataset, which involves per-prompt checklist items and item-by-item judging.

Training Details

The training involved 522 prompts and 14,616 learner pairs over 2 pipeline epochs. Key parameters included a learning rate of 3e-7, AdamW optimizer, and a maximum sequence length of 2048 tokens. Training signal was generated by a local open-weight judge, averaging 10 judgments per pair.

Evaluation and Limitations

Evaluation was conducted using greedy generation against released baseline answers, judged by a local open-weight model. On Arena-Hard, it scored 0.6111, and on AlpacaEval, it achieved 0.4118. It's crucial to note that these scores are not directly comparable to public leaderboards, as a local open-weight judge was used instead of standard GPT-4 based evaluators. The differences between this arm and its counterparts (NBPO and PROSPER) are not statistically resolved, indicating that the fixed-reference Nash approach did not significantly outperform or underperform the other aggregation rules in this specific experimental setup.

Key Characteristics

  • Fixed-reference Nash aggregation: Uses a frozen reference policy with per-prompt multipliers.
  • WildChecklists training: Optimized for scenarios where prompts have specific checklist items for evaluation.
  • Comparative research focus: Designed to study the impact of different aggregation rules on model performance.