C12H18/gats-coldstart-sft-grpo-qwen2.5-7b-webshop
C12H18/gats-coldstart-sft-grpo-qwen2.5-7b-webshop is a 7.6 billion parameter language model, part of the Cold-start collection for the GATS paper on capability-aware on-policy distillation. This model is a Qwen2.5-7B-Instruct student, fine-tuned with Supervised Fine-Tuning (SFT) cold-start followed by pure GRPO (Grouped Reinforcement Learning with Policy Optimization) using a Qwen2.5-1.5B teacher on WebShop tasks. It is specifically intended for research into cold-start strategies for agent reinforcement learning and reproduction of GATS paper baselines, achieving a final evaluation score of 75.61 ± 3.44.
Loading preview...
Overview
This model, C12H18/gats-coldstart-sft-grpo-qwen2.5-7b-webshop, is a 7.6 billion parameter Qwen2.5-7B-Instruct student model. It is a component of the Cold-start collection developed for the GATS paper, focusing on capability-aware on-policy distillation. The model's training regimen involved an initial Supervised Fine-Tuning (SFT) cold-start phase, followed by pure GRPO (Grouped Reinforcement Learning with Policy Optimization). A Qwen2.5-1.5B model, pre-trained with GRPO on WebShop train tasks, served as the teacher.
Key Training Details
- Student Model: Qwen2.5-7B-Instruct
- Teacher Model: Qwen2.5-1.5B GRPO-trained on WebShop
- Training Process: SFT cold-start (3 epochs, lr 1e-5) followed by 150 steps of pure GRPO (lr 1e-6, gamma 0.95, KL loss coef 0.01).
- Performance: Achieved a final evaluation score of 75.61 ± 3.44 on WebShop tasks.
Intended Use
This model is specifically designed for:
- Research into cold-start strategies for agent reinforcement learning (e.g., comparing imitation-first vs. RL-from-scratch approaches).
- Reproduction and validation of baselines presented in the GATS paper.
It is not intended as a general-purpose assistant but rather a specialized tool for RL research.