davidheineman/opd-teacher-Q2.5I-CampsitePuzzle-step149
The davidheineman/opd-teacher-Q2.5I-CampsitePuzzle-step149 is a 1.5 billion parameter Qwen2.5-1.5B-Instruct model, developed by davidheineman, specifically trained as a teacher model using Reinforcement Learning from Vision-based Environments (RLVE). It was fine-tuned with GRPO for 150 updates on the CampsitePuzzle environment at difficulty 0, making it specialized for on-policy distillation experiments within this specific environment. This model is designed to provide expert guidance for agents learning to solve the CampsitePuzzle.
Loading preview...
Model Overview
This model, opd-teacher-Q2.5I-CampsitePuzzle-step149, is a specialized teacher model based on the Qwen2.5-1.5B-Instruct architecture. Developed by davidheineman, it features 1.5 billion parameters and was trained using Reinforcement Learning from Vision-based Environments (RLVE).
Key Training Details
- Base Model: Qwen/Qwen2.5-1.5B-Instruct
- Training Method: Trained with GRPO (Generalized Reinforcement Policy Optimization) for 150 updates.
- Environment: Specifically fine-tuned on the
CampsitePuzzleenvironment at difficulty 0. - Purpose: Intended for use in 32-environment on-policy distillation experiments, serving as an expert teacher.
- Checkpoint:
step149represents the final, 150th update checkpoint, with weights converted to Hugging Face safetensors.
Intended Use
This model is designed to provide expert demonstrations or guidance within the CampsitePuzzle environment for agents undergoing on-policy distillation. Its specialized training makes it highly effective for tasks related to this specific environment and experimental setup.