nics-efc/MARSHAL-Generalist-Qwen3-4B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Nov 28, 2025License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

MARSHAL-Generalist-Qwen3-4B is a 4 billion parameter language model developed by the MARSHAL framework team, initialized from Qwen3-4B. It is specifically trained via self-play on diverse strategic games like Tic-Tac-Toe, Kuhn Poker, and Mini Hanabi, encompassing competitive, cooperative, perfect, and imperfect information settings. This model excels at multi-agent reasoning, demonstrating significant performance improvements on held-out games and consistent gains in leading multi-agent systems on reasoning benchmarks, making it suitable for complex strategic interactions.

Loading preview...

MARSHAL-Generalist-Qwen3-4B Overview

This model is the generalist variant of the MARSHAL framework, built upon Qwen3-4B with 4 billion parameters and a 32K context length. It has been uniquely trained using an end-to-end reinforcement learning framework that incentivizes multi-agent reasoning through self-play in a diverse array of strategic games. These games include competitive (Tic-Tac-Toe, Kuhn Poker) and cooperative (Mini Hanabi) scenarios, covering both perfect and imperfect information settings.

Key Differentiators & Capabilities

  • Multi-Agent Reasoning: MARSHAL addresses credit assignment challenges in multi-agent, multi-turn self-play through a turn-level advantage estimator and agent-specific advantage normalization, ensuring accurate learning signals.
  • Strategic Game Mastery: Demonstrates strong performance in various strategic games, achieving up to 28.7% improvement on held-out games not seen during training.
  • Generalization to Reasoning Benchmarks: When integrated into multi-agent systems (MASs), MARSHAL consistently improves performance on reasoning benchmarks, showing gains of up to +10.0% on AIME and +7.6% on GPQA-Diamond.

Ideal Use Cases

This model is particularly well-suited for applications requiring advanced multi-agent reasoning and strategic decision-making. Developers can leverage MARSHAL-Generalist-Qwen3-4B for:

  • Developing AI agents for complex competitive and cooperative games.
  • Enhancing existing multi-agent systems with improved reasoning capabilities.
  • Research into multi-agent reinforcement learning and strategic AI.