davidheineman/opd-teacher-Q2.5I-LinkBeads-step149
This is a davidheineman/opd-teacher-Q2.5I-LinkBeads-step149 model, a 1.5 billion parameter Qwen2.5-1.5B-Instruct teacher model. It was trained using RLVE with GRPO for 150 updates, specifically optimized for the LinkBeads environment at difficulty 0. This model is intended for on-policy distillation experiments within a 32-environment setup, demonstrating specialized performance in its target environment.
Loading preview...
Model Overview
This model, davidheineman/opd-teacher-Q2.5I-LinkBeads-step149, is a specialized 1.5 billion parameter instruction-tuned teacher model based on the Qwen2.5-1.5B-Instruct architecture. It has been specifically trained using Reinforcement Learning from Value Estimates (RLVE) with the GRPO algorithm over 150 updates.
Key Capabilities
- Environment Specialization: Optimized for the
LinkBeadsenvironment at difficulty 0. - Training Methodology: Utilizes GRPO (Generalized Reinforcement Policy Optimization) for training, indicating a focus on robust policy learning.
- Distillation Target: Designed as a teacher model for on-policy distillation experiments across 32 environments.
- Context Length: Supports a context length of 32768 tokens.
Intended Use Cases
This model is primarily intended for research and development in the field of on-policy distillation, particularly for scenarios involving the LinkBeads environment. Its specialized training makes it suitable for acting as a teacher in a student-teacher learning setup, where a smaller student model learns from the teacher's policy.