davidheineman/opd-teacher-Q2.5I-CampsitePuzzle-step149

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 28, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The davidheineman/opd-teacher-Q2.5I-CampsitePuzzle-step149 is a 1.5 billion parameter Qwen2.5-1.5B-Instruct model, developed by davidheineman, specifically trained as a teacher model using Reinforcement Learning from Vision-based Environments (RLVE). It was fine-tuned with GRPO for 150 updates on the CampsitePuzzle environment at difficulty 0, making it specialized for on-policy distillation experiments within this specific environment. This model is designed to provide expert guidance for agents learning to solve the CampsitePuzzle.

Loading preview...

Model Overview

This model, opd-teacher-Q2.5I-CampsitePuzzle-step149, is a specialized teacher model based on the Qwen2.5-1.5B-Instruct architecture. Developed by davidheineman, it features 1.5 billion parameters and was trained using Reinforcement Learning from Vision-based Environments (RLVE).

Key Training Details

  • Base Model: Qwen/Qwen2.5-1.5B-Instruct
  • Training Method: Trained with GRPO (Generalized Reinforcement Policy Optimization) for 150 updates.
  • Environment: Specifically fine-tuned on the CampsitePuzzle environment at difficulty 0.
  • Purpose: Intended for use in 32-environment on-policy distillation experiments, serving as an expert teacher.
  • Checkpoint: step149 represents the final, 150th update checkpoint, with weights converted to Hugging Face safetensors.

Intended Use

This model is designed to provide expert demonstrations or guidance within the CampsitePuzzle environment for agents undergoing on-policy distillation. Its specialized training makes it highly effective for tasks related to this specific environment and experimental setup.