launch/MET-D-Qwen3-4B-en-only
MET-D-Qwen3-4B-en-only is an English-only moral reasoning model fine-tuned from Qwen3-4B. It is designed to judge the acceptability of actions within moral dilemmas from a specific character's perspective, providing a chain-of-thought explanation. The model answers two questions: whether an action is acceptable and if it would cause emotional discomfort. It is specifically trained on self-generated, rejection-sampled reasoning traces in English.
Loading preview...
Overview
MET-D-Qwen3-4B-en-only is a specialized English-only moral reasoning model, fine-tuned from the Qwen3-4B base model. It is designed to analyze moral dilemmas by taking into account a character's description and a candidate action. The model generates a judgment from the character's perspective, providing an explicit chain-of-thought explanation before delivering its final answers.
Key Capabilities
- Moral Judgment: Evaluates the acceptability of a given action within a moral dilemma from a specified character's viewpoint.
- Emotional Discomfort Assessment: Determines if performing (or not performing) an action would cause mental or emotional discomfort for the character.
- Chain-of-Thought Reasoning: Provides a detailed reasoning trace for its judgments, enhancing transparency and interpretability.
- English-Only Processing: This specific variant is trained and optimized exclusively for English language input and output.
Training and Uniqueness
The model was trained using self-generated reasoning traces, which were then rejection-sampled against ground-truth answers derived from specific character perspectives. This method addresses the complexity of verifying moral reasoning by grounding it in a defined character's theoretical framework. It is part of the broader MET collection which includes multilingual and other single-language variants based on different base models like Qwen3-4B, Qwen3-8B, and Gemma-3-4B-it.