davidheineman/opd-teacher-Q2.5I-CantorExpansion-step149
The davidheineman/opd-teacher-Q2.5I-CantorExpansion-step149 is a 1.5 billion parameter Qwen2.5-Instruct model, developed by davidheineman, specifically trained as a teacher model using Reinforcement Learning from Vision-based Environments (RLVE). It was fine-tuned on the CantorExpansion environment at difficulty 0 for an on-policy distillation experiment. This model is optimized for demonstrating learned behaviors within a specific, controlled environment, showcasing capabilities in RLVE-driven instruction.
Loading preview...
Model Overview
This model, opd-teacher-Q2.5I-CantorExpansion-step149, is a specialized 1.5 billion parameter Qwen2.5-Instruct teacher model. It was developed by davidheineman and trained using Reinforcement Learning from Vision-based Environments (RLVE) on the CantorExpansion environment at difficulty 0.
Key Capabilities
- RLVE-Trained Teacher: Functions as a teacher model within an on-policy distillation experiment, specifically for the
CantorExpansionenvironment. - Qwen2.5-1.5B-Instruct Base: Built upon the robust Qwen2.5-1.5B-Instruct architecture, providing a strong foundation for instruction-following.
- Targeted Training: Underwent 150 updates with GRPO, focusing its learning on specific environmental interactions.
Good For
- On-Policy Distillation Research: Ideal for researchers and developers working on on-policy distillation experiments, particularly those involving the
CantorExpansionenvironment. - RLVE Demonstrations: Useful for understanding and demonstrating the application of RLVE in training teacher models for specific tasks.
- Specialized Instruction Following: Provides a case study for how base instruction models can be fine-tuned for highly specific, environment-driven instructional roles.