idealab-cs2/reappraisal-4b-grpo-committee
The idealab-cs2/reappraisal-4b-grpo-committee model is a 4 billion parameter GRPO fine-tune of Qwen/Qwen3-4B-Thinking-2507, developed by idealab-cs2. It is specifically designed to generate effective cognitive reappraisals for negative interpersonal scenarios, aiming to match or exceed GPT-4-0314's performance. This model's unique differentiator is its training against a two-member committee of independent reward scorers, ensuring high-quality, agreed-upon reappraisals in two sentences or fewer.
Loading preview...
Model Overview
The idealab-cs2/reappraisal-4b-grpo-committee model is a specialized 4 billion parameter language model, fine-tuned from Qwen/Qwen3-4B-Thinking-2507 using Group Relative Policy Optimization (GRPO). Its primary function is to generate concise, two-sentence cognitive reappraisals for negative interpersonal scenarios, aiming to reduce negative emotions.
Key Differentiators
- Committee-Based Reward Training: Unlike models trained with a single reward model, this policy is trained against a unique two-member committee of reward scorers. This committee comprises a discriminative regression reward model and a generative reward model, with rewards only applied when both scorers agree. This approach, using an equal-weight (0.5/0.5) combination with a disagreement guard, helps prevent the policy from gaming a single reward function.
- Targeted Reappraisal Generation: The model is specifically optimized for cognitive reframing, producing alternative interpretations of situations addressed to the person in the scenario.
- Performance: On the 6 vignettes it was trained on, the model demonstrates a win-rate of 0.87 against GPT-4-0314 and 0.67 against DeepSeek-R1-671B in pairwise evaluations. However, off-distribution, its performance aligns more with its base model capacity, indicating its specialized nature.
Intended Use & Limitations
This model is a research artifact for studying RLHF, reward modeling, and computational emotion regulation. It is designed to write short cognitive reappraisals of negative interpersonal situations. It is not intended for clinical or mental-health applications and its quality is validated only on this narrow reappraisal task.