C12H18/gats-coldstart-opd-grpo-qwen2.5-7b-sciworld

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 29, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

C12H18/gats-coldstart-opd-grpo-qwen2.5-7b-sciworld is a 7.6 billion parameter Qwen2.5-7B-Instruct model, fine-tuned using a pure-OPD cold-start followed by GRPO (Grouped Reinforcement Learning with Policy Optimization) on SciWorld tasks. This model is specifically designed for research into cold-start strategies for agent reinforcement learning, focusing on imitation-first versus RL-from-scratch approaches. It achieves an evaluation score of 37.76 ± 2.51 on seen tasks and 33.07 ± 3.93 on unseen tasks, making it suitable for reproducing baselines of the GATS paper.

Loading preview...

GATS Cold-Start: Pure-OPD Cold-Start + GRPO (SciWorld)

This model, C12H18/gats-coldstart-opd-grpo-qwen2.5-7b-sciworld, is a 7.6 billion parameter Qwen2.5-7B-Instruct student model. It is part of the Cold-start collection developed for the GATS paper, which focuses on capability-aware on-policy distillation.

Key Training Methodology

The model's training involved a two-stage process:

  • Pure-OPD Cold-Start: Initialized from a pure-OPD cold-start checkpoint (C12H18/gats-coldstart-opd-qwen2.5-7b-sciworld), which involved 29 on-policy distillation steps.
  • Pure GRPO Fine-tuning: Subsequently, it underwent 150 steps of pure GRPO (Grouped Reinforcement Learning with Policy Optimization) with specific hyperparameters, including a learning rate of 1e-6, group size n=8, and a KL loss coefficient of 0.01. The teacher model for this process was a Qwen2.5-1.5B model, GRPO-trained on SciWorld tasks.

Performance

Upon final evaluation, the model achieved:

  • Seen tasks: 37.76 ± 2.51
  • Unseen tasks: 33.07 ± 3.93 (evaluated with K=3 and rollout seeds 0/1/2)

Intended Use Cases

This model is specifically intended for:

  • Research on cold-start strategies for agent reinforcement learning, particularly comparing imitation-first versus RL-from-scratch approaches.
  • Reproduction of baselines presented in the GATS paper.

It is not designed as a general-purpose assistant.