davidheineman/opd-teacher-Q2.5I-DeltaMinPopcount-step149

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 28, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The davidheineman/opd-teacher-Q2.5I-DeltaMinPopcount-step149 is a 1.5 billion parameter Qwen2.5-1.5B-Instruct model, fine-tuned by davidheineman. This model was specifically trained using Reinforcement Learning from VLM Environments (RLVE) on the DeltaMinPopcount environment. It is designed for on-policy distillation experiments within a 32-environment setup, focusing on learning optimal policies for specific tasks.

Loading preview...

Model Overview

This model, opd-teacher-Q2.5I-DeltaMinPopcount-step149, is a specialized 1.5 billion parameter Qwen2.5-1.5B-Instruct variant developed by davidheineman. It has been fine-tuned using Reinforcement Learning from VLM Environments (RLVE) with the GRPO algorithm over 150 updates.

Key Characteristics

  • Base Model: Built upon Qwen/Qwen2.5-1.5B-Instruct.
  • Training Method: Utilizes Reinforcement Learning from VLM Environments (RLVE) with GRPO.
  • Environment Specificity: Trained on the DeltaMinPopcount environment at difficulty 0.
  • Purpose: Intended for use in a 32-environment on-policy distillation experiment.
  • Checkpoint: Represents the final, 150th update (step 149) of the training process.

Intended Use Cases

This model is specifically designed for research and experimentation in:

  • On-Policy Distillation: Serving as a teacher model for distilling policies in complex environments.
  • Reinforcement Learning Research: Investigating RLVE and GRPO training methodologies.
  • Environment-Specific Task Learning: Exploring model performance on the DeltaMinPopcount environment.