davidheineman/opd-teacher-Q2.5I-FractionalProgramming-step149
TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 28, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold
The davidheineman/opd-teacher-Q2.5I-FractionalProgramming-step149 is a 1.5 billion parameter Qwen2.5-1.5B-Instruct teacher model, developed by davidheineman. It was trained using RLVE with GRPO for 150 updates on the FractionalProgramming environment. This model is specifically designed for on-policy distillation experiments within a 32-environment setup, focusing on fractional programming tasks.
Loading preview...
Model Overview
This model, opd-teacher-Q2.5I-FractionalProgramming-step149, is a specialized Qwen2.5-1.5B-Instruct teacher model developed by davidheineman. It features 1.5 billion parameters and was trained using Reinforcement Learning from Value Estimates (RLVE) with the GRPO algorithm.
Key Characteristics
- Base Model: Built upon Qwen/Qwen2.5-1.5B-Instruct.
- Training: Underwent 150 updates using GRPO, specifically on the
FractionalProgrammingenvironment at difficulty 0. - Context Length: Supports a context length of 32768 tokens.
- Purpose: Primarily intended for use in a 32-environment on-policy distillation experiment, focusing on teaching fractional programming concepts.
Use Cases
This model is particularly suited for:
- Research in RLVE and On-Policy Distillation: Ideal for experiments involving teacher models in reinforcement learning setups.
- Fractional Programming Environments: Designed to provide guidance or act as a teacher within environments centered around fractional programming tasks.
- Specialized AI Training: Useful for developers and researchers working on fine-tuning or distilling knowledge for specific, narrow AI domains.