3l3ktr4/donorsim-qwen3-8b-modeAB-step51
The 3l3ktr4/donorsim-qwen3-8b-modeAB-step51 is an 8 billion parameter Qwen3 model fine-tuned by 3l3ktr4 using GRPO on the iterated Donor's Game. This model specializes in simulating complex social interactions within group dynamics, specifically focusing on scenarios with mixed real groupmate discussions. It is designed for research into cooperative behaviors and strategic decision-making in game theory contexts.
Loading preview...
Model Overview
This model, donorsim-qwen3-8b-modeAB-step51, is an 8 billion parameter Qwen3 variant developed by 3l3ktr4. It has been meticulously fine-tuned using the GRPO (verl 0.7.1) method, incorporating LoRA (r16/alpha32) merged into bf16 weights. The primary training focus is on simulating the iterated Donor's Game, specifically Stage 2, which involves group games with varying group sizes (K in {2,4,6}) and within-group partner rotation.
Key Capabilities
- Specialized Game Theory Simulation: Optimized for modeling complex interactions in the iterated Donor's Game.
- Mixed Mode A/B Scenarios: Trained on scenarios where 50% include real groupmate discussion, enhancing its ability to simulate nuanced social dynamics.
- Iterative Training: Represents the 51st step in a continuous training lineage, building upon previous stages with fixed and rotating partners.
- Direct Loadability: The full weights are merged, allowing direct loading with
transformersor vLLM without requiring an adapter.
Good For
- Research in Social Simulation: Ideal for academics and researchers studying cooperative behavior, strategic decision-making, and game theory.
- Understanding Group Dynamics: Useful for analyzing how LLMs can model and predict outcomes in multi-agent interaction scenarios.
- Exploring GRPO Fine-tuning: Provides an example of a Qwen3 model fine-tuned with GRPO for a specific, complex task.