C12H18/gats-coldstart-sft-grpo-qwen2.5-7b-alfworld
C12H18/gats-coldstart-sft-grpo-qwen2.5-7b-alfworld is a 7.6 billion parameter Qwen2.5-7B-Instruct model, part of the GATS paper's Cold-Start collection. This model is fine-tuned using Supervised Fine-Tuning (SFT) and Grouped Reinforcement Learning from Human Feedback (GRPO) on ALFWorld tasks. It is specifically designed for research into cold-start strategies for agent reinforcement learning and reproducing GATS paper baselines, rather than general-purpose assistance.
Loading preview...
Model Overview
This model, C12H18/gats-coldstart-sft-grpo-qwen2.5-7b-alfworld, is a 7.6 billion parameter Qwen2.5-7B-Instruct student model. It is a component of the GATS paper's Cold-Start collection, focusing on capability-aware on-policy distillation. The model undergoes a two-phase training process: an SFT cold-start followed by pure GRPO.
Key Training Details
- Student Model: Qwen2.5-7B-Instruct
- Teacher Model: Qwen2.5-1.5B, pre-trained with GRPO on ALFWorld train tasks.
- SFT Cold-Start Phase: Trained for 3 epochs with a learning rate of 1e-5 on a corpus of 13,970 rows / 239 tasks, using a multiturn format.
- GRPO Phase: Applied for 150 steps with specific hyperparameters (e.g., lr 1e-6, gamma 0.95, KL loss coefficient 0.01).
Performance
- Final Evaluation (K=3, rollout seeds 0/1/2):
- In-Distribution (ID): 77.86 ± 1.63
- Out-of-Distribution (OOD): 83.85 ± 6.74
Intended Use Cases
This model is specifically intended for:
- Research on cold-start strategies for agent reinforcement learning (e.g., comparing imitation-first vs. RL-from-scratch approaches).
- Reproduction of baselines presented in the GATS paper.
It is not designed as a general-purpose assistant.