YiPz/qwen3-4b-pokerbench-grpo

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jan 15, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

YiPz/qwen3-4b-pokerbench-grpo is a 4 billion parameter Qwen3-based causal language model developed by YiPz, specifically fine-tuned for optimal poker decision-making. This model leverages Group Relative Policy Optimization (GRPO) and LLM-as-judge rewards to refine its strategic outputs. It excels at providing structured reasoning and precise poker actions, making it suitable for advanced poker analysis and simulation.

Loading preview...

Model Overview

This model, YiPz/qwen3-4b-pokerbench-grpo, is a 4 billion parameter Qwen3-based language model specifically fine-tuned for poker decision-making. It was developed by YiPz using a two-stage training process to enhance its strategic capabilities in poker scenarios.

Key Capabilities

  • Specialized Poker Strategy: The model is optimized to analyze poker situations and provide strategic actions, including structured reasoning for its decisions.
  • Reinforcement Learning Refinement: It underwent a Group Relative Policy Optimization (GRPO) stage, where it was further refined using LLM-as-judge rewards to reinforce better poker decisions.
  • Structured Output: The model generates outputs with a clear reasoning section (<think>) followed by a specific action (<action>), facilitating easy integration and interpretation.
  • Base Model: Built upon the Qwen/Qwen3-4B-thinking-2507 base model, initially fine-tuned on high-quality reasoning traces from PokerBench.

Good For

  • Poker Analysis: Ideal for applications requiring deep analysis of poker hands and strategic recommendations.
  • Poker AI Development: Useful for developers building poker bots or AI agents that need to make informed decisions.
  • Simulation and Training: Can be employed in poker simulations or as a training tool to understand optimal play.