3l3ktr4/donorsim-qwen3-8b-abstract-step50
The 3l3ktr4/donorsim-qwen3-8b-abstract-step50 is an 8 billion parameter Qwen3 model fine-tuned by 3l3ktr4 using GRPO for simulating iterated Donor's Game scenarios. It processes naturalistic descriptions of social interactions and outputs choices, focusing on abstract everyday scenes rather than explicit game theory terms. This model is specialized for research into social dynamics and decision-making within simulated environments, particularly concerning reciprocity and partner memory.
Loading preview...
Model Overview
This model, donorsim-qwen3-8b-abstract-step50, is an 8 billion parameter Qwen3 variant fine-tuned by 3l3ktr4. It utilizes GRPO (verl 0.7.1) with LoRA (r16/alpha32) merged into bf16 weights. The primary focus of its fine-tuning is the abstract / everyday-scene stage of the iterated Donor's Game.
Key Capabilities and Training
- Simulated Social Interaction: The model is designed to interpret short, naturalistic descriptions of social situations involving named individuals in small groups.
- Decision Output: It generates
CHOICE: 1orCHOICE: 2responses, with option order randomized per turn. - Reward Structure: Training incorporates rewards based on normalized payoff (Term 1) and reciprocity (Term 2), without explicit group or CFE terms.
- Partner Dynamics: It simulates partner rotation within a roster and includes per-partner memory. Re-encounter probability (
w) and gossip probability (q) are described in words each turn. - Iterative Fine-tuning: This version represents 50 abstract steps, building upon
donorsim-qwen3-8b-modeAB-step75, which involved 75 steps of a structured group game.
Usage
The model's full weights are merged, allowing direct loading with transformers or vLLM without requiring an adapter. It is suitable for research applications focused on understanding and simulating complex social decision-making processes in abstract, natural language contexts.