idealab-cs2/reappraisal-4b-grpo-committee

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 7, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The idealab-cs2/reappraisal-4b-grpo-committee model is a 4 billion parameter GRPO fine-tune of Qwen/Qwen3-4B-Thinking-2507, developed by ruggsea. It is specifically optimized for generating cognitive reappraisals of negative interpersonal scenarios, aiming to match or exceed GPT-4-0314's effectiveness. This model uniquely employs a two-member committee of different-mechanism reward scorers to ensure high-quality, effective reappraisals in two sentences or fewer.

Loading preview...

Model Overview

This model, Reappraisal-4B-GRPO-Committee, is a 4 billion parameter language model developed by ruggsea, fine-tuned from Qwen/Qwen3-4B-Thinking-2507 using Group Relative Policy Optimization (GRPO). Its primary function is to generate short, effective cognitive reappraisals for negative interpersonal scenarios, aiming to reinterpret situations to reduce negative emotions.

Key Differentiators & Capabilities

  • Specialized Reappraisal Generation: Trained on scenarios from Li et al. (2025), it produces alternative interpretations in two sentences or fewer, addressed to the person in the scenario.
  • Committee-Based Reward System: Utilizes a novel two-member committee of reward scorers (a discriminative regression model and a generative preference model) hosted at idealab-cs2/reappraisal-reward-model-v2. This committee approach, rewarding only where both scorers agree, enhances robustness and prevents gaming of individual reward models.
  • Strong Performance on Task: Achieves a win-rate of 0.872 against GPT-4-0314 on the 6 training vignettes, and even out-reappraises DeepSeek-R1-671B pairwise on these specific scenarios.

Intended Use & Limitations

This model is a research artifact for studying RLHF and computational emotion regulation. It is not intended for clinical or mental-health applications. While highly effective on its narrow training task, its generalization to out-of-distribution scenarios is limited, performing at parity or below strong open 72B models like Qwen2.5-72B and losing to DeepSeek-R1-671B on fresh scenarios.