davidheineman/opd-teacher-Q2.5I-LinkBeads-step149

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 28, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

This is a davidheineman/opd-teacher-Q2.5I-LinkBeads-step149 model, a 1.5 billion parameter Qwen2.5-1.5B-Instruct teacher model. It was trained using RLVE with GRPO for 150 updates, specifically optimized for the LinkBeads environment at difficulty 0. This model is intended for on-policy distillation experiments within a 32-environment setup, demonstrating specialized performance in its target environment.

Loading preview...

Model Overview

This model, davidheineman/opd-teacher-Q2.5I-LinkBeads-step149, is a specialized 1.5 billion parameter instruction-tuned teacher model based on the Qwen2.5-1.5B-Instruct architecture. It has been specifically trained using Reinforcement Learning from Value Estimates (RLVE) with the GRPO algorithm over 150 updates.

Key Capabilities

  • Environment Specialization: Optimized for the LinkBeads environment at difficulty 0.
  • Training Methodology: Utilizes GRPO (Generalized Reinforcement Policy Optimization) for training, indicating a focus on robust policy learning.
  • Distillation Target: Designed as a teacher model for on-policy distillation experiments across 32 environments.
  • Context Length: Supports a context length of 32768 tokens.

Intended Use Cases

This model is primarily intended for research and development in the field of on-policy distillation, particularly for scenarios involving the LinkBeads environment. Its specialized training makes it suitable for acting as a teacher in a student-teacher learning setup, where a smaller student model learns from the teacher's policy.