davidheineman/opd-teacher-Q2.5I-GoldWashing-step149

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 28, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

davidheineman/opd-teacher-Q2.5I-GoldWashing-step149 is a 1.5 billion parameter Qwen2.5-1.5B-Instruct model, developed by davidheineman. This model is a teacher agent trained with Reinforcement Learning from Value Estimates (RLVE) using GRPO on the 'GoldWashing' environment at difficulty 0. It is specifically designed for on-policy distillation experiments within a 32-environment setup, serving as a specialized component for RL research.

Loading preview...

Model Overview

This model, opd-teacher-Q2.5I-GoldWashing-step149, is a specialized 1.5 billion parameter Qwen2.5-1.5B-Instruct teacher model. Developed by davidheineman, it was trained using Reinforcement Learning from Value Estimates (RLVE) with the GRPO algorithm. The training focused on the GoldWashing environment at difficulty 0, completing 150 updates, with step149 representing the final checkpoint.

Key Capabilities

  • RLVE Teacher Model: Functions as a teacher in a Reinforcement Learning from Value Estimates (RLVE) setup.
  • Environment Specificity: Optimized for the GoldWashing environment at difficulty 0.
  • On-Policy Distillation: Intended for use in 32-environment on-policy distillation experiments.
  • GRPO Training: Trained using the GRPO algorithm over 150 updates.

Technical Details

Intended Use

This model is primarily for researchers and developers working on reinforcement learning, particularly in the context of on-policy distillation and teacher-student learning paradigms within controlled environments like GoldWashing.