idealab-cs2/reappraisal-4b-grpo-rmv2
Reappraisal-4B-GRPO-RMv2 by idealab-cs2 is a 4 billion parameter model, fine-tuned from Qwen/Qwen3-4B-Thinking-2507 using Group Relative Policy Optimization (GRPO). This model specializes in generating cognitive reappraisals for negative interpersonal scenarios, aiming to reinterpret situations to reduce negative emotions. It is specifically optimized to produce effective, short reappraisals, outperforming GPT-4-0314 on its training vignettes.
Loading preview...
Overview
Reappraisal-4B-GRPO-RMv2 is a specialized 4 billion parameter language model developed by idealab-cs2, fine-tuned from Qwen/Qwen3-4B-Thinking-2507. Its core function is to generate cognitive reappraisals, which are alternative interpretations of negative interpersonal scenarios designed to reduce associated negative emotions. The model was trained using Group Relative Policy Optimization (GRPO) with a reward model (idealab-cs2/reappraisal-reward-model-v2) derived from human effectiveness ratings.
Key Capabilities
- Cognitive Reappraisal Generation: Produces short, two-sentence reappraisals for negative interpersonal situations.
- Targeted Optimization: Specifically trained to match or exceed the effectiveness of GPT-4-0314's reappraisals on its training vignettes.
- GRPO Fine-tuning: Utilizes Group Relative Policy Optimization for robust policy learning.
Evaluation and Limitations
Evaluated against GPT-4-0314, this model achieved a win-rate of 0.806 on the 6 specific interpersonal vignettes it was trained on, as judged by Llama-3.1-70B-Instruct. It's crucial to note that this model is a specialist trained on 6 scenarios and does not generalize past parity on out-of-distribution scenarios, losing to models like DeepSeek-R1-671B off-distribution. Its performance is validated only on this narrow task and should not be used for clinical or mental-health applications. The improvement over previous versions is largely attributed to the enhanced reward model.