davidheineman/opd-teacher-Q2.5I-DeltaMinPopcount-step149
The davidheineman/opd-teacher-Q2.5I-DeltaMinPopcount-step149 is a 1.5 billion parameter Qwen2.5-1.5B-Instruct model, fine-tuned by davidheineman. This model was specifically trained using Reinforcement Learning from VLM Environments (RLVE) on the DeltaMinPopcount environment. It is designed for on-policy distillation experiments within a 32-environment setup, focusing on learning optimal policies for specific tasks.
Loading preview...
Model Overview
This model, opd-teacher-Q2.5I-DeltaMinPopcount-step149, is a specialized 1.5 billion parameter Qwen2.5-1.5B-Instruct variant developed by davidheineman. It has been fine-tuned using Reinforcement Learning from VLM Environments (RLVE) with the GRPO algorithm over 150 updates.
Key Characteristics
- Base Model: Built upon Qwen/Qwen2.5-1.5B-Instruct.
- Training Method: Utilizes Reinforcement Learning from VLM Environments (RLVE) with GRPO.
- Environment Specificity: Trained on the
DeltaMinPopcountenvironment at difficulty 0. - Purpose: Intended for use in a 32-environment on-policy distillation experiment.
- Checkpoint: Represents the final, 150th update (step 149) of the training process.
Intended Use Cases
This model is specifically designed for research and experimentation in:
- On-Policy Distillation: Serving as a teacher model for distilling policies in complex environments.
- Reinforcement Learning Research: Investigating RLVE and GRPO training methodologies.
- Environment-Specific Task Learning: Exploring model performance on the
DeltaMinPopcountenvironment.