davidheineman/opd-teacher-Q2.5I-GoldWashing-step149
davidheineman/opd-teacher-Q2.5I-GoldWashing-step149 is a 1.5 billion parameter Qwen2.5-1.5B-Instruct model, developed by davidheineman. This model is a teacher agent trained with Reinforcement Learning from Value Estimates (RLVE) using GRPO on the 'GoldWashing' environment at difficulty 0. It is specifically designed for on-policy distillation experiments within a 32-environment setup, serving as a specialized component for RL research.
Loading preview...
Model Overview
This model, opd-teacher-Q2.5I-GoldWashing-step149, is a specialized 1.5 billion parameter Qwen2.5-1.5B-Instruct teacher model. Developed by davidheineman, it was trained using Reinforcement Learning from Value Estimates (RLVE) with the GRPO algorithm. The training focused on the GoldWashing environment at difficulty 0, completing 150 updates, with step149 representing the final checkpoint.
Key Capabilities
- RLVE Teacher Model: Functions as a teacher in a Reinforcement Learning from Value Estimates (RLVE) setup.
- Environment Specificity: Optimized for the
GoldWashingenvironment at difficulty 0. - On-Policy Distillation: Intended for use in 32-environment on-policy distillation experiments.
- GRPO Training: Trained using the GRPO algorithm over 150 updates.
Technical Details
- Base Model: Built upon Qwen/Qwen2.5-1.5B-Instruct.
- Training Run: Details available via W&B 16beff02.
- Training Code: Utilizes the davidheineman/rlve repository.
Intended Use
This model is primarily for researchers and developers working on reinforcement learning, particularly in the context of on-policy distillation and teacher-student learning paradigms within controlled environments like GoldWashing.