davidheineman/opd-teacher-Q2.5I-MafMafia-step149
The davidheineman/opd-teacher-Q2.5I-MafMafia-step149 is a 1.5 billion parameter Qwen2.5-Instruct model, fine-tuned using Reinforcement Learning from Value Equivalents (RLVE). It was specifically trained on the 'MafMafia' environment at difficulty 0 for an on-policy distillation experiment. This model is optimized for performance within the specified environment, demonstrating specialized agent behavior.
Loading preview...
Model Overview
This model, opd-teacher-Q2.5I-MafMafia-step149, is a 1.5 billion parameter Qwen2.5-Instruct teacher model developed by davidheineman. It has been specifically fine-tuned using Reinforcement Learning from Value Equivalents (RLVE) within the MafMafia environment at difficulty 0.
Key Capabilities
- Specialized Agent Behavior: Trained to act as a teacher model for on-policy distillation experiments, demonstrating learned behaviors within the
MafMafiaenvironment. - Reinforcement Learning Fine-tuning: Utilizes GRPO (Generalized Policy Optimization) for 150 updates, indicating a focus on robust policy learning.
- Base Model: Built upon the
Qwen/Qwen2.5-1.5B-Instructarchitecture, providing a strong foundation for instruction-following and language understanding.
Intended Use
This model is primarily intended for research and experimentation related to on-policy distillation, particularly within the context of the 32-environment experiment it was designed for. Its specialized training makes it suitable for analyzing agent performance and transfer learning within the MafMafia environment.