davidheineman/opd-teacher-Q2.5I-Cornfield-step149

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 28, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The davidheineman/opd-teacher-Q2.5I-Cornfield-step149 model is a 1.5 billion parameter Qwen2.5-1.5B-Instruct teacher model, developed by davidheineman. It was trained using Reinforcement Learning from Vision-based Environments (RLVE) on the Cornfield environment with a 32768 token context length. This model is specifically designed for on-policy distillation experiments, serving as a teacher for agent training in simulated environments.

Loading preview...

Model Overview

This model, opd-teacher-Q2.5I-Cornfield-step149, is a specialized Qwen2.5-1.5B-Instruct teacher model with 1.5 billion parameters and a 32768 token context length. Developed by davidheineman, it is specifically trained for on-policy distillation (OPD) experiments within a 32-environment setup.

Key Characteristics

  • Base Model: Built upon Qwen/Qwen2.5-1.5B-Instruct.
  • Training Method: Trained using Reinforcement Learning from Vision-based Environments (RLVE), specifically with GRPO (Generalized Policy Optimization).
  • Environment: Optimized for the Cornfield environment at difficulty 0.
  • Training Duration: Underwent 150 updates, with step149 representing the final checkpoint.
  • Purpose: Intended to act as a teacher model for distilling policies in reinforcement learning scenarios.

Use Cases

This model is particularly suited for:

  • Reinforcement Learning Research: Specifically for experiments involving on-policy distillation.
  • Agent Training: Serving as a teacher to guide the learning of student agents in simulated environments like Cornfield.
  • Policy Transfer: Investigating methods for transferring learned policies from a teacher model to a smaller or different student model.