davidheineman/opd-teacher-Q2.5I-Disinfection-step149
The davidheineman/opd-teacher-Q2.5I-Disinfection-step149 is a 1.5 billion parameter Qwen2.5-Instruct model developed by davidheineman. It is specifically trained using RLVE on the 'Disinfection' environment at difficulty 0, making it an on-policy distillation teacher model. This model is optimized for specific reinforcement learning environments, distinguishing it from general-purpose LLMs. It features a context length of 32768 tokens and is intended for specialized experimental setups.
Loading preview...
Model Overview
The davidheineman/opd-teacher-Q2.5I-Disinfection-step149 is a specialized 1.5 billion parameter instruction-tuned model based on the Qwen2.5 architecture. Developed by davidheineman, this model functions as an on-policy distillation (OPD) teacher, specifically trained within a reinforcement learning environment.
Key Characteristics
- Base Model: Utilizes the
Qwen/Qwen2.5-1.5B-Instructas its foundation. - Training Method: Trained using Reinforcement Learning from Value Equivalents (RLVE) with the GRPO algorithm.
- Environment Specificity: Fine-tuned exclusively on the
Disinfectionenvironment at difficulty 0. - Training Duration: Underwent 150 updates, with
step149representing the final checkpoint. - Context Length: Supports a substantial context window of 32768 tokens.
Intended Use
This model is designed for experimental purposes, particularly within the context of a 32-environment on-policy distillation experiment. Its specialized training makes it suitable for research and development in reinforcement learning and model distillation within defined environments, rather than general conversational or creative tasks.