C12H18/gats-coldstart-sft-grpo-qwen2.5-7b-alfworld

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 29, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

C12H18/gats-coldstart-sft-grpo-qwen2.5-7b-alfworld is a 7.6 billion parameter Qwen2.5-7B-Instruct model, part of the GATS paper's Cold-Start collection. This model is fine-tuned using Supervised Fine-Tuning (SFT) and Grouped Reinforcement Learning from Human Feedback (GRPO) on ALFWorld tasks. It is specifically designed for research into cold-start strategies for agent reinforcement learning and reproducing GATS paper baselines, rather than general-purpose assistance.

Loading preview...

Model Overview

This model, C12H18/gats-coldstart-sft-grpo-qwen2.5-7b-alfworld, is a 7.6 billion parameter Qwen2.5-7B-Instruct student model. It is a component of the GATS paper's Cold-Start collection, focusing on capability-aware on-policy distillation. The model undergoes a two-phase training process: an SFT cold-start followed by pure GRPO.

Key Training Details

  • Student Model: Qwen2.5-7B-Instruct
  • Teacher Model: Qwen2.5-1.5B, pre-trained with GRPO on ALFWorld train tasks.
  • SFT Cold-Start Phase: Trained for 3 epochs with a learning rate of 1e-5 on a corpus of 13,970 rows / 239 tasks, using a multiturn format.
  • GRPO Phase: Applied for 150 steps with specific hyperparameters (e.g., lr 1e-6, gamma 0.95, KL loss coefficient 0.01).

Performance

  • Final Evaluation (K=3, rollout seeds 0/1/2):
    • In-Distribution (ID): 77.86 ± 1.63
    • Out-of-Distribution (OOD): 83.85 ± 6.74

Intended Use Cases

This model is specifically intended for:

  • Research on cold-start strategies for agent reinforcement learning (e.g., comparing imitation-first vs. RL-from-scratch approaches).
  • Reproduction of baselines presented in the GATS paper.

It is not designed as a general-purpose assistant.