davidheineman/opd-teacher-Q2.5I-Cornfield-step149
The davidheineman/opd-teacher-Q2.5I-Cornfield-step149 model is a 1.5 billion parameter Qwen2.5-1.5B-Instruct teacher model, developed by davidheineman. It was trained using Reinforcement Learning from Vision-based Environments (RLVE) on the Cornfield environment with a 32768 token context length. This model is specifically designed for on-policy distillation experiments, serving as a teacher for agent training in simulated environments.
Loading preview...
Model Overview
This model, opd-teacher-Q2.5I-Cornfield-step149, is a specialized Qwen2.5-1.5B-Instruct teacher model with 1.5 billion parameters and a 32768 token context length. Developed by davidheineman, it is specifically trained for on-policy distillation (OPD) experiments within a 32-environment setup.
Key Characteristics
- Base Model: Built upon Qwen/Qwen2.5-1.5B-Instruct.
- Training Method: Trained using Reinforcement Learning from Vision-based Environments (RLVE), specifically with GRPO (Generalized Policy Optimization).
- Environment: Optimized for the
Cornfieldenvironment at difficulty 0. - Training Duration: Underwent 150 updates, with
step149representing the final checkpoint. - Purpose: Intended to act as a teacher model for distilling policies in reinforcement learning scenarios.
Use Cases
This model is particularly suited for:
- Reinforcement Learning Research: Specifically for experiments involving on-policy distillation.
- Agent Training: Serving as a teacher to guide the learning of student agents in simulated environments like
Cornfield. - Policy Transfer: Investigating methods for transferring learned policies from a teacher model to a smaller or different student model.