davidheineman/opd-teacher-Q2.5I-ConvexHull-step149
The davidheineman/opd-teacher-Q2.5I-ConvexHull-step149 is a 1.5 billion parameter Qwen2.5-Instruct model, developed by davidheineman, specifically fine-tuned using Reinforcement Learning from Vision-based Environments (RLVE). This model is trained on the ConvexHull environment at difficulty 0, making it specialized for on-policy distillation experiments within this specific visual environment. It is intended as a teacher model for tasks requiring environmental interaction and learning from visual cues.
Loading preview...
Model Overview
This model, opd-teacher-Q2.5I-ConvexHull-step149, is a specialized Qwen2.5-1.5B-Instruct teacher model developed by davidheineman. It has been fine-tuned using Reinforcement Learning from Vision-based Environments (RLVE), specifically targeting the ConvexHull environment at difficulty 0. The training involved 150 updates using the GRPO algorithm, with step149 representing the final checkpoint.
Key Capabilities and Training
- Base Model: Built upon the robust Qwen/Qwen2.5-1.5B-Instruct architecture.
- Specialized Fine-tuning: Underwent RLVE training, indicating its proficiency in learning from and interacting with visual environments.
- Environment Specificity: Explicitly trained on the
ConvexHullenvironment, suggesting optimized performance for tasks within this domain. - Distillation Focus: Designed as a "teacher" model for on-policy distillation experiments, implying its role in transferring learned policies to other models.
Intended Use Cases
This model is particularly well-suited for:
- Research in RLVE: Ideal for experiments involving reinforcement learning in vision-based environments, especially those related to
ConvexHull. - On-Policy Distillation: Serves as a teacher model for distilling learned policies to student models.
- Environmental Interaction: Useful for tasks requiring an understanding and interaction within specific visual environments at a foundational difficulty level.