3l3ktr4/donorsim-qwen3-8b-abstract-step40

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 29, 2026Architecture:Transformer Featherless Exclusive Cold

The 3l3ktr4/donorsim-qwen3-8b-abstract-step40 is an 8 billion parameter Qwen3 model, fine-tuned with GRPO on an iterated Donor's Game. This model specializes in interpreting naturalistic social situations and making binary choices based on normalized payoff and reciprocity, rather than explicit game theory terms. It is designed to simulate decision-making in abstract, everyday social scenarios with rotating partners and memory of past interactions.

Loading preview...

Overview

3l3ktr4/donorsim-qwen3-8b-abstract-step40 is an 8 billion parameter Qwen3 model that has undergone specialized fine-tuning using the GRPO (verl 0.7.1) method. This model is distinctively trained on the "abstract / everyday-scene" stage of an iterated Donor's Game, focusing on naturalistic social interactions rather than explicit game-theoretic constructs.

Key Capabilities

  • Social Decision-Making: Interprets short, naturalistic descriptions of social situations involving named individuals within a small group.
  • Binary Choice Output: Generates CHOICE: 1 or CHOICE: 2 responses, with randomized option order, based on the perceived social context.
  • Reciprocity-Based Rewards: Training incorporates rewards based on normalized payoff and reciprocity, without using explicit "cooperate" or "defect" vocabulary.
  • Contextual Memory: Manages per-partner memory within a roster of rotating players, influencing decisions based on past interactions.
  • Dynamic Social Parameters: Accounts for re-encounter probability (w) and gossip probability (q), phrased in natural language for each turn.

Training Details

This model builds upon donorsim-qwen3-8b-modeAB-step75 with an additional 40 abstract steps. The fine-tuning involved merging LoRA adapters (r16/alpha32) into bf16 weights, making it directly loadable with transformers or vLLM without requiring separate adapters.