C12H18/gats-coldstart-opd-grpo-qwen2.5-7b-sciworld
C12H18/gats-coldstart-opd-grpo-qwen2.5-7b-sciworld is a 7.6 billion parameter Qwen2.5-7B-Instruct model, fine-tuned using a pure-OPD cold-start followed by GRPO (Grouped Reinforcement Learning with Policy Optimization) on SciWorld tasks. This model is specifically designed for research into cold-start strategies for agent reinforcement learning, focusing on imitation-first versus RL-from-scratch approaches. It achieves an evaluation score of 37.76 ± 2.51 on seen tasks and 33.07 ± 3.93 on unseen tasks, making it suitable for reproducing baselines of the GATS paper.
Loading preview...
GATS Cold-Start: Pure-OPD Cold-Start + GRPO (SciWorld)
This model, C12H18/gats-coldstart-opd-grpo-qwen2.5-7b-sciworld, is a 7.6 billion parameter Qwen2.5-7B-Instruct student model. It is part of the Cold-start collection developed for the GATS paper, which focuses on capability-aware on-policy distillation.
Key Training Methodology
The model's training involved a two-stage process:
- Pure-OPD Cold-Start: Initialized from a pure-OPD cold-start checkpoint (
C12H18/gats-coldstart-opd-qwen2.5-7b-sciworld), which involved 29 on-policy distillation steps. - Pure GRPO Fine-tuning: Subsequently, it underwent 150 steps of pure GRPO (Grouped Reinforcement Learning with Policy Optimization) with specific hyperparameters, including a learning rate of 1e-6, group size n=8, and a KL loss coefficient of 0.01. The teacher model for this process was a Qwen2.5-1.5B model, GRPO-trained on SciWorld tasks.
Performance
Upon final evaluation, the model achieved:
- Seen tasks: 37.76 ± 2.51
- Unseen tasks: 33.07 ± 3.93 (evaluated with K=3 and rollout seeds 0/1/2)
Intended Use Cases
This model is specifically intended for:
- Research on cold-start strategies for agent reinforcement learning, particularly comparing imitation-first versus RL-from-scratch approaches.
- Reproduction of baselines presented in the GATS paper.
It is not designed as a general-purpose assistant.