C12H18/gats-coldstart-sft-grpo-qwen2.5-7b-webshop

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 29, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

C12H18/gats-coldstart-sft-grpo-qwen2.5-7b-webshop is a 7.6 billion parameter language model, part of the Cold-start collection for the GATS paper on capability-aware on-policy distillation. This model is a Qwen2.5-7B-Instruct student, fine-tuned with Supervised Fine-Tuning (SFT) cold-start followed by pure GRPO (Grouped Reinforcement Learning with Policy Optimization) using a Qwen2.5-1.5B teacher on WebShop tasks. It is specifically intended for research into cold-start strategies for agent reinforcement learning and reproduction of GATS paper baselines, achieving a final evaluation score of 75.61 ± 3.44.

Loading preview...

Overview

This model, C12H18/gats-coldstart-sft-grpo-qwen2.5-7b-webshop, is a 7.6 billion parameter Qwen2.5-7B-Instruct student model. It is a component of the Cold-start collection developed for the GATS paper, focusing on capability-aware on-policy distillation. The model's training regimen involved an initial Supervised Fine-Tuning (SFT) cold-start phase, followed by pure GRPO (Grouped Reinforcement Learning with Policy Optimization). A Qwen2.5-1.5B model, pre-trained with GRPO on WebShop train tasks, served as the teacher.

Key Training Details

  • Student Model: Qwen2.5-7B-Instruct
  • Teacher Model: Qwen2.5-1.5B GRPO-trained on WebShop
  • Training Process: SFT cold-start (3 epochs, lr 1e-5) followed by 150 steps of pure GRPO (lr 1e-6, gamma 0.95, KL loss coef 0.01).
  • Performance: Achieved a final evaluation score of 75.61 ± 3.44 on WebShop tasks.

Intended Use

This model is specifically designed for:

  • Research into cold-start strategies for agent reinforcement learning (e.g., comparing imitation-first vs. RL-from-scratch approaches).
  • Reproduction and validation of baselines presented in the GATS paper.

It is not intended as a general-purpose assistant but rather a specialized tool for RL research.