XingYing-stack/TIPS-Qwen3-4B-Instruct-2507-Agent

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 27, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The TIPS-Qwen3-4B-Instruct-2507-Agent is a 4 billion parameter generative reward model developed by XingYing-stack, initialized from Qwen/Qwen3-4B-Instruct-2507. It is specifically trained with outcome-only GRPO for step-level verification and first-error localization within multi-turn agent trajectories. This model excels at process verification and reward modeling, rather than general-purpose conversational tasks, and supports a 32768 token context length.

Loading preview...

Model Overview

The TIPS-Qwen3-4B-Instruct-2507-Agent is a specialized 4 billion parameter model developed by XingYing-stack. It is built upon the Qwen/Qwen3-4B-Instruct-2507 base and has undergone further training using outcome-only Generative Reward Policy Optimization (GRPO). This model is designed as a generative reward model, focusing on evaluating and verifying steps within complex multi-turn agent interactions.

Key Capabilities

  • Step-level Verification: Assesses the correctness and quality of individual steps in an agent's operational sequence.
  • First-Error Localization: Identifies the precise point where an error first occurs within a multi-turn agent trajectory.
  • Generative Reward Modeling: Provides feedback or scores based on the outcomes of agent actions, facilitating process improvement.
  • Specialized Training: Optimized specifically for agent trajectory analysis, distinguishing it from general-purpose instruction-tuned models.

Intended Use Cases

This model is not intended for general-purpose chat or conversational AI. Its primary applications include:

  • Reward Modeling: Generating rewards or feedback signals for reinforcement learning in agent systems.
  • Process Verification: Validating the execution flow and correctness of automated agent tasks.
  • Agent Trajectory Analysis: Debugging and understanding failures in complex multi-step agent operations.

Training data and code for this model are publicly available, allowing for further research and application development. Users should refer to the associated TIPS repository for specific prompt templates and evaluation scripts to ensure proper utilization.