RealPirate786/Minesweeper_agent_Qwen3_4B_GRPO

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 22, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The RealPirate786/Minesweeper_agent_Qwen3_4B_GRPO is a 4 billion parameter Qwen3 model developed by RealPirate786. It was fine-tuned from unsloth/qwen3-4b-unsloth-bnb-4bit using Unsloth and Huggingface's TRL library, enabling 2x faster training. This model is specifically designed as an agent for Minesweeper, leveraging its fine-tuning for strategic gameplay.

Loading preview...

Model Overview

The RealPirate786/Minesweeper_agent_Qwen3_4B_GRPO is a 4 billion parameter Qwen3 model developed by RealPirate786. It has been fine-tuned from the unsloth/qwen3-4b-unsloth-bnb-4bit base model. The fine-tuning process utilized Unsloth and Huggingface's TRL library, which facilitated a 2x speedup in training.

Key Capabilities

  • Minesweeper Agent: This model is specifically fine-tuned to act as an agent for playing Minesweeper, indicating its specialization in strategic decision-making within this game context.
  • Efficient Training: Leverages Unsloth for accelerated fine-tuning, making the development process more efficient.
  • Qwen3 Architecture: Built upon the Qwen3 architecture, providing a robust foundation for its capabilities.

Good For

  • Minesweeper Game AI: Ideal for research and development related to AI agents for the game Minesweeper.
  • Exploring Fine-tuning Techniques: Demonstrates the application of Unsloth and TRL for efficient model adaptation.
  • Small-scale Agent Development: Suitable for projects requiring a specialized agent with a 4 billion parameter footprint.