C12H18/gats-coldstart-opd-qwen2.5-7b-webshop
C12H18/gats-coldstart-opd-qwen2.5-7b-webshop is a 7.6 billion parameter Qwen2.5-7B-Instruct model, developed by C12H18, fine-tuned for on-policy distillation (OPD) within the GATS framework. It is specifically trained to imitate a Qwen2.5-1.5B teacher model on WebShop tasks, focusing on cold-start strategies for agent reinforcement learning. This model is intended for research into imitation-first versus RL-from-scratch approaches and reproducing GATS paper baselines.
Loading preview...
Model Overview
C12H18/gats-coldstart-opd-qwen2.5-7b-webshop is a specialized 7.6 billion parameter model based on Qwen2.5-7B-Instruct, developed by C12H18. It is part of the GATS (capability-aware on-policy distillation) cold-start collection, designed for research into agent reinforcement learning.
Key Capabilities & Training
- On-Policy Distillation (OPD): The model undergoes a pure imitation objective, rolling out on-policy in an environment and imitating a Qwen2.5-1.5B teacher model. The teacher was GRPO-trained on WebShop train tasks.
- Cold-Start Phase: This specific model represents the cold-start phase, where the student model performs 29 steps of pure imitation, with a validation score of 60.2% before any RL training.
- Training Configuration: Rollouts involved 16 tasks per step, with groups of 8, using a learning rate of 1e-6.
Intended Use Cases
- Research on Cold-Start Strategies: Ideal for investigating different cold-start approaches in agent RL, such as imitation-first versus RL-from-scratch.
- GATS Paper Reproduction: Suitable for reproducing baselines and experiments described in the GATS paper.
Limitations
- Not a General-Purpose Assistant: This model is highly specialized for its research purpose and is not intended for general conversational or assistant tasks.