davidheineman/opd-teacher-Q2.5I-ConvexHull-step149

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 28, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The davidheineman/opd-teacher-Q2.5I-ConvexHull-step149 is a 1.5 billion parameter Qwen2.5-Instruct model, developed by davidheineman, specifically fine-tuned using Reinforcement Learning from Vision-based Environments (RLVE). This model is trained on the ConvexHull environment at difficulty 0, making it specialized for on-policy distillation experiments within this specific visual environment. It is intended as a teacher model for tasks requiring environmental interaction and learning from visual cues.

Loading preview...

Model Overview

This model, opd-teacher-Q2.5I-ConvexHull-step149, is a specialized Qwen2.5-1.5B-Instruct teacher model developed by davidheineman. It has been fine-tuned using Reinforcement Learning from Vision-based Environments (RLVE), specifically targeting the ConvexHull environment at difficulty 0. The training involved 150 updates using the GRPO algorithm, with step149 representing the final checkpoint.

Key Capabilities and Training

  • Base Model: Built upon the robust Qwen/Qwen2.5-1.5B-Instruct architecture.
  • Specialized Fine-tuning: Underwent RLVE training, indicating its proficiency in learning from and interacting with visual environments.
  • Environment Specificity: Explicitly trained on the ConvexHull environment, suggesting optimized performance for tasks within this domain.
  • Distillation Focus: Designed as a "teacher" model for on-policy distillation experiments, implying its role in transferring learned policies to other models.

Intended Use Cases

This model is particularly well-suited for:

  • Research in RLVE: Ideal for experiments involving reinforcement learning in vision-based environments, especially those related to ConvexHull.
  • On-Policy Distillation: Serves as a teacher model for distilling learned policies to student models.
  • Environmental Interaction: Useful for tasks requiring an understanding and interaction within specific visual environments at a foundational difficulty level.