C12H18/gats-coldstart-opd-qwen2.5-7b-alfworld

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 29, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The C12H18/gats-coldstart-opd-qwen2.5-7b-alfworld model is a 7.6 billion parameter Qwen2.5-7B-Instruct student model, part of the GATS paper's Cold-start collection for capability-aware on-policy distillation. It is specifically trained for research into cold-start strategies for agent reinforcement learning, imitating a Qwen2.5-1.5B teacher model on ALFWorld tasks. This model is optimized for agent RL research, focusing on imitation-first learning in simulated environments.

Loading preview...

GATS Cold-Start: Pure-OPD for ALFWorld

This model, C12H18/gats-coldstart-opd-qwen2.5-7b-alfworld, is a 7.6 billion parameter Qwen2.5-7B-Instruct student model developed by C12H18. It is a component of the GATS (capability-aware on-policy distillation) paper's Cold-start collection, specifically designed for research into agent reinforcement learning (RL) cold-start strategies.

Key Characteristics

  • Architecture: Based on the Qwen2.5-7B-Instruct model.
  • Training Method: Utilizes pure on-policy distillation (OPD), where the student model imitates a Qwen2.5-1.5B teacher model that was GRPO-trained on ALFWorld tasks.
  • Cold-Start Phase: Trained for 52 steps with a pure imitation objective (opd_only coefficient 1.0), without an RL reward term.
  • Performance: Achieved a validation score of 57.8% on the ALFWorld out-of-distribution split before any RL steps.

Intended Use Cases

  • Research: Primarily intended for research on cold-start strategies in agent RL, comparing imitation-first approaches versus RL-from-scratch.
  • GATS Paper Reproduction: Useful for reproducing baselines presented in the GATS paper.
  • Specialized Agent Training: Not designed as a general-purpose assistant, but rather for specialized tasks within simulated environments like ALFWorld.