xiamoent/Agent-G2-webshop-7b

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 2, 2026Architecture:Transformer Featherless Exclusive Cold

The xiamoent/Agent-G2-webshop-7b is a 7.6 billion parameter language agent checkpoint developed by xiamoent, initialized from Qwen2.5-7B-Instruct. It is specifically post-trained with Agent-G2: Gaussian Guidance for Agentic Reinforcement Learning, making it highly specialized for tasks within the WebShop simulator. This model excels at long-horizon language agent research and agentic reinforcement learning, achieving a 92.3 reward score and 84.4% final-purchase success in the WebShop environment.

Loading preview...

Agent-G2 WebShop 7B Overview

Agent-G2 WebShop 7B is a specialized language agent model, fine-tuned from Qwen2.5-7B-Instruct with 7.6 billion parameters and a 32,768 token context length. Developed by xiamoent, this checkpoint leverages Agent-G2: Gaussian Guidance for Agentic Reinforcement Learning to enhance its performance in complex, multi-step environments.

Key Capabilities & Features

  • WebShop Specialization: Optimized for the sandboxed WebShop e-commerce simulator, demonstrating strong performance in product search and purchase tasks.
  • Gaussian Guidance: Utilizes an adaptive Gaussian distribution to sample expert-prefix depth during training, improving policy optimization without additional probe rollouts.
  • Agentic Reinforcement Learning: Designed for long-horizon tasks, enabling the model to reason and select actions effectively within an environment.
  • High Performance: Achieves a 92.3 Reward Score and 84.4% Final-purchase Success on the WebShop benchmark.

Intended Use Cases

  • Reproducing Agent-G2 Results: Ideal for researchers aiming to validate findings from the Agent-G2 project.
  • Research on Language Agents: Suitable for studies on long-horizon language agents and agentic reinforcement learning methodologies.
  • Adaptive Expert-Prefix Guidance: Useful for investigating and evaluating adaptive expert-prefix guidance techniques in agent training.
  • Action Selection Evaluation: Designed for evaluating action selection mechanisms over environment-provided admissible action sets within the WebShop environment.