davidheineman/opd-teacher-Q2.5I-MafMafia-step149

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 28, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The davidheineman/opd-teacher-Q2.5I-MafMafia-step149 is a 1.5 billion parameter Qwen2.5-Instruct model, fine-tuned using Reinforcement Learning from Value Equivalents (RLVE). It was specifically trained on the 'MafMafia' environment at difficulty 0 for an on-policy distillation experiment. This model is optimized for performance within the specified environment, demonstrating specialized agent behavior.

Loading preview...

Model Overview

This model, opd-teacher-Q2.5I-MafMafia-step149, is a 1.5 billion parameter Qwen2.5-Instruct teacher model developed by davidheineman. It has been specifically fine-tuned using Reinforcement Learning from Value Equivalents (RLVE) within the MafMafia environment at difficulty 0.

Key Capabilities

  • Specialized Agent Behavior: Trained to act as a teacher model for on-policy distillation experiments, demonstrating learned behaviors within the MafMafia environment.
  • Reinforcement Learning Fine-tuning: Utilizes GRPO (Generalized Policy Optimization) for 150 updates, indicating a focus on robust policy learning.
  • Base Model: Built upon the Qwen/Qwen2.5-1.5B-Instruct architecture, providing a strong foundation for instruction-following and language understanding.

Intended Use

This model is primarily intended for research and experimentation related to on-policy distillation, particularly within the context of the 32-environment experiment it was designed for. Its specialized training makes it suitable for analyzing agent performance and transfer learning within the MafMafia environment.