XingYing-stack/TIPS-Qwen3-4B-Thinking-2507-Agent

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 27, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The TIPS-Qwen3-4B-Thinking-2507-Agent is a 4 billion parameter generative reward model developed by XingYing-stack, initialized from Qwen3-4B-Thinking-2507. It is specifically trained with outcome-only GRPO for step-level verification and first-error localization within multi-turn agent trajectories. This model excels at process verification and reward modeling, rather than general-purpose conversational tasks.

Loading preview...

Model Overview

The TIPS-Qwen3-4B-Thinking-2507-Agent is a specialized 4 billion parameter model developed by XingYing-stack. It is initialized from the Qwen/Qwen3-4B-Thinking-2507 base model and has undergone further training using an outcome-only Generative Reward Policy Optimization (GRPO) approach.

Key Capabilities

  • Generative Reward Modeling: Functions as a reward model to evaluate and provide feedback on agent actions.
  • Step-Level Verification: Capable of verifying individual steps within complex multi-turn agent trajectories.
  • First-Error Localization: Designed to identify and pinpoint the initial error in a sequence of agent operations.

Intended Use Cases

This model is primarily intended for:

  • Process Verification: Assessing the correctness and efficiency of agent workflows.
  • Reward Modeling: Generating rewards for reinforcement learning setups involving agent trajectories.

It is important to note that this checkpoint is not designed for general-purpose chat or conversational AI. Users should refer to the TIPS repository for specific prompt templates and evaluation scripts tailored to its intended use. Training data and code are available on Hugging Face Datasets.