davidheineman/opd-teacher-Q2.5I-Differentiate-step149

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 28, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The davidheineman/opd-teacher-Q2.5I-Differentiate-step149 is a 1.5 billion parameter Qwen2.5-1.5B-Instruct model, fine-tuned by davidheineman using Reinforcement Learning from Value Equivalents (RLVE). This model is specifically trained on the 'Differentiate' environment at difficulty 0, making it specialized for tasks related to differentiation. It is intended for on-policy distillation experiments, demonstrating a focused application in a specific RL environment.

Loading preview...

Model Overview

This model, davidheineman/opd-teacher-Q2.5I-Differentiate-step149, is a specialized instruction-tuned language model based on the Qwen2.5-1.5B-Instruct architecture. Developed by davidheineman, it features 1.5 billion parameters and a context length of 32768 tokens.

Key Capabilities and Training

The primary distinction of this model lies in its training methodology and specific application:

  • Reinforcement Learning from Value Equivalents (RLVE): The model was trained using RLVE with GRPO for 150 updates, indicating a focus on learning optimal policies within a defined environment.
  • Specialized Environment Training: It is specifically trained on the Differentiate environment at difficulty 0. This suggests a strong proficiency in tasks related to differentiation, likely within a simulated or structured problem-solving context.
  • Experimental Purpose: The model is part of a larger 32-environment on-policy distillation experiment, highlighting its role in research and development for RL-driven language model applications.

Intended Use Case

This model is particularly suited for:

  • Research in Reinforcement Learning: Ideal for researchers exploring on-policy distillation, RLVE, and GRPO training methods.
  • Differentiation-related Tasks: Its specialized training on the Differentiate environment makes it a strong candidate for tasks requiring understanding or execution of differentiation principles.
  • Experimental Prototyping: Useful for developers and researchers working on similar controlled environments or seeking to understand the impact of specific RL training paradigms on language models.