davidheineman/opd-teacher-Q2.5I-EmperorWorries-step149

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 28, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

davidheineman/opd-teacher-Q2.5I-EmperorWorries-step149 is a 1.5 billion parameter Qwen2.5-1.5B-Instruct teacher model developed by davidheineman. It was trained using Reinforcement Learning from Value Equivalents (RLVE) with GRPO for 150 updates on the 'EmperorWorries' environment. This model is specifically designed for on-policy distillation experiments within a 32-environment setup, leveraging its 32768 token context length.

Loading preview...

Model Overview

davidheineman/opd-teacher-Q2.5I-EmperorWorries-step149 is a specialized 1.5 billion parameter instruction-tuned model based on the Qwen2.5-1.5B-Instruct architecture. Developed by davidheineman, this model serves as a 'teacher' in a Reinforcement Learning from Value Equivalents (RLVE) framework.

Key Characteristics

  • Base Model: Built upon Qwen/Qwen2.5-1.5B-Instruct.
  • Training Method: Trained using GRPO (Generalized Reinforcement Policy Optimization) within the RLVE paradigm.
  • Environment Specificity: Fine-tuned on the EmperorWorries environment at difficulty 0.
  • Training Duration: Underwent 150 updates, with step149 representing the final checkpoint.
  • Purpose: Intended for use in 32-environment on-policy distillation experiments.
  • Context Length: Supports a substantial context window of 32768 tokens.

Intended Use Cases

This model is specifically engineered for research and development in reinforcement learning, particularly for:

  • On-Policy Distillation: Acting as a teacher model to guide the training of student policies.
  • RLVE Experiments: Exploring and validating concepts related to Reinforcement Learning from Value Equivalents.
  • Environment-Specific Tasks: Performing within the EmperorWorries environment, leveraging its specialized training.