davidheineman/opd-teacher-Q2.5I-Sorting-step149

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 28, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

davidheineman/opd-teacher-Q2.5I-Sorting-step149 is a 1.5 billion parameter Qwen2.5-Instruct model, fine-tuned by davidheineman using Reinforcement Learning from VLM Feedback (RLVE). This model is specifically trained as a 'teacher' for the 'Sorting' environment at difficulty 0, intended for on-policy distillation experiments. It specializes in generating optimal sorting strategies within this constrained environment.

Loading preview...

Model Overview

davidheineman/opd-teacher-Q2.5I-Sorting-step149 is a specialized 1.5 billion parameter language model based on the Qwen2.5-1.5B-Instruct architecture. It has been fine-tuned by davidheineman using Reinforcement Learning from VLM Feedback (RLVE) specifically for the Sorting environment at difficulty 0.

Key Characteristics

  • Base Model: Qwen2.5-1.5B-Instruct, providing a strong foundation for instruction following.
  • Training Method: Utilizes GRPO (Generalized Reinforcement Policy Optimization) for 150 updates, focusing on generating optimal actions within a specific task.
  • Specialization: Functions as a 'teacher' model, designed to provide expert guidance or demonstrations for the Sorting environment.
  • Context Length: Supports a context length of 32768 tokens, allowing for processing of relatively long sequences related to sorting tasks.

Intended Use

This model is primarily intended for research in on-policy distillation experiments, where it serves as a source of expert behavior for a 'student' model learning to perform sorting tasks. Its training is highly specific to the Sorting environment at difficulty 0, making it suitable for controlled experimental setups in reinforcement learning and policy transfer.