C12H18/gats-coldstart-opd-qwen2.5-7b-alfworld
The C12H18/gats-coldstart-opd-qwen2.5-7b-alfworld model is a 7.6 billion parameter Qwen2.5-7B-Instruct student model, part of the GATS paper's Cold-start collection for capability-aware on-policy distillation. It is specifically trained for research into cold-start strategies for agent reinforcement learning, imitating a Qwen2.5-1.5B teacher model on ALFWorld tasks. This model is optimized for agent RL research, focusing on imitation-first learning in simulated environments.
Loading preview...
GATS Cold-Start: Pure-OPD for ALFWorld
This model, C12H18/gats-coldstart-opd-qwen2.5-7b-alfworld, is a 7.6 billion parameter Qwen2.5-7B-Instruct student model developed by C12H18. It is a component of the GATS (capability-aware on-policy distillation) paper's Cold-start collection, specifically designed for research into agent reinforcement learning (RL) cold-start strategies.
Key Characteristics
- Architecture: Based on the Qwen2.5-7B-Instruct model.
- Training Method: Utilizes pure on-policy distillation (OPD), where the student model imitates a Qwen2.5-1.5B teacher model that was GRPO-trained on ALFWorld tasks.
- Cold-Start Phase: Trained for 52 steps with a pure imitation objective (opd_only coefficient 1.0), without an RL reward term.
- Performance: Achieved a validation score of 57.8% on the ALFWorld out-of-distribution split before any RL steps.
Intended Use Cases
- Research: Primarily intended for research on cold-start strategies in agent RL, comparing imitation-first approaches versus RL-from-scratch.
- GATS Paper Reproduction: Useful for reproducing baselines presented in the GATS paper.
- Specialized Agent Training: Not designed as a general-purpose assistant, but rather for specialized tasks within simulated environments like ALFWorld.