hongli-zhan/MINT-empathy-Qwen3-1.7B
MINT-empathy-Qwen3-1.7B is a 1.7 billion parameter Qwen3-based model fine-tuned by Hongli Zhan for multi-turn empathic dialogue. Utilizing the MINT (Multi-turn Inter-tactic Novelty Training) reinforcement learning framework, this checkpoint optimizes for empathic response quality. It achieves a significant improvement in aggregate empathy, making it suitable for research in supportive response generation and discourse diversity.
Loading preview...
MINT-empathy-Qwen3-1.7B Overview
This model, developed by Hongli Zhan, is a 1.7 billion parameter variant of the Qwen3 architecture, specifically fine-tuned for generating empathic responses in multi-turn dialogues. It leverages the MINT (Multi-turn Inter-tactic Novelty Training) reinforcement learning framework, which optimizes both empathy quality and cross-turn discourse-move novelty.
Key Capabilities & Performance
- Enhanced Empathy: Demonstrates a substantial improvement in aggregate empathy, increasing from 3.60 to 4.54 on the Lend-an-Ear test set compared to the vanilla Qwen3-1.7B baseline.
- Reinforcement Learning: Trained using GRPO via VERL, incorporating rewards for empathy quality and cross-turn tactic diversity.
- Research Focus: Primarily intended as a research artifact for studying empathic dialogue, discourse diversity, and supportive response generation.
Intended Use Cases
- Empathic Dialogue Research: Ideal for academic and research applications exploring the generation of empathetic conversational agents.
- Discourse Diversity Studies: Useful for investigating and improving the variety of conversational tactics in AI responses.
- Supportive Response Generation: Can be applied in scenarios requiring AI to provide more understanding and supportive interactions.
It's important to note that at this 1.7B scale, the primary improvement is in empathy quality, with less significant reduction in tactic repetition compared to larger MINT models.