ThakiCloud/RAG-Gate-8B
ThakiCloud/RAG-Gate-8B is an 8 billion parameter model, fine-tuned from Qwen/Qwen3-8B, designed to act as a decision-making gate in Retrieval-Augmented Generation (RAG) pipelines. It processes a question and retrieved passages to determine whether to 'Answer', 'Retrieve' more evidence, or 'Stop'. This model significantly improves action accuracy to 0.949, reducing unsupported answers and over-refusal rates, making it ideal for enhancing the reliability and efficiency of RAG systems by intelligently managing the flow of information.
Loading preview...
RAG-Gate-8B Overview
ThakiCloud/RAG-Gate-8B is an 8 billion parameter model, fine-tuned from Qwen/Qwen3-8B, specifically engineered to operate as a critical decision-making component within Retrieval-Augmented Generation (RAG) pipelines. Positioned after retrieval and before generation, its primary function is to evaluate a given question, the retrieved passages, and the availability of further retrieval. Based on this assessment, it outputs one of three actions: Answer (if sufficient evidence is present), Retrieve (if more evidence is needed and available), or Stop (if insufficient evidence and no further retrieval is possible).
Key Capabilities
- Enhanced RAG Decision-Making: Achieves an action accuracy of 0.949 on a blind test set, a substantial improvement over the base model's 0.529.
- Reduced Unsupported Answers: Significantly lowers the rate of answering without sufficient support to 0.047 (from 0.103).
- Minimized Over-Refusal: Drastically decreases instances where the model refuses to answer despite having sufficient evidence, down to 0.057 (from 0.606).
- Robustness to Distractors: Maintains high accuracy (0.962) even when retrieved evidence includes edited distractor passages (
FULL_DECOYstate). - Contextual Judgment: Judges sufficiency based on the completeness of the support chain, rather than surface edits or prior knowledge.
Good For
- Optimizing RAG Workflows: Ideal for developers looking to build more reliable and efficient RAG systems by automating the decision of when to answer, retrieve, or stop.
- Improving Answer Quality: Ensures that answers are only generated when supported by a complete chain of evidence, reducing hallucinations and unsupported claims.
- Resource Management: Intelligently manages retrieval calls, preventing unnecessary computations when evidence is already sufficient or when no further retrieval would help.
- Complex Question Answering: Particularly effective in multi-hop question answering scenarios where evidence sufficiency is critical.