ThakiCloud/RAG-Gate-8B

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 5, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

ThakiCloud/RAG-Gate-8B is an 8 billion parameter model, fine-tuned from Qwen/Qwen3-8B, designed to act as a decision-making gate in Retrieval-Augmented Generation (RAG) pipelines. It processes a question and retrieved passages to determine whether to 'Answer', 'Retrieve' more evidence, or 'Stop'. This model significantly improves action accuracy to 0.949, reducing unsupported answers and over-refusal rates, making it ideal for enhancing the reliability and efficiency of RAG systems by intelligently managing the flow of information.

Loading preview...

RAG-Gate-8B Overview

ThakiCloud/RAG-Gate-8B is an 8 billion parameter model, fine-tuned from Qwen/Qwen3-8B, specifically engineered to operate as a critical decision-making component within Retrieval-Augmented Generation (RAG) pipelines. Positioned after retrieval and before generation, its primary function is to evaluate a given question, the retrieved passages, and the availability of further retrieval. Based on this assessment, it outputs one of three actions: Answer (if sufficient evidence is present), Retrieve (if more evidence is needed and available), or Stop (if insufficient evidence and no further retrieval is possible).

Key Capabilities

  • Enhanced RAG Decision-Making: Achieves an action accuracy of 0.949 on a blind test set, a substantial improvement over the base model's 0.529.
  • Reduced Unsupported Answers: Significantly lowers the rate of answering without sufficient support to 0.047 (from 0.103).
  • Minimized Over-Refusal: Drastically decreases instances where the model refuses to answer despite having sufficient evidence, down to 0.057 (from 0.606).
  • Robustness to Distractors: Maintains high accuracy (0.962) even when retrieved evidence includes edited distractor passages (FULL_DECOY state).
  • Contextual Judgment: Judges sufficiency based on the completeness of the support chain, rather than surface edits or prior knowledge.

Good For

  • Optimizing RAG Workflows: Ideal for developers looking to build more reliable and efficient RAG systems by automating the decision of when to answer, retrieve, or stop.
  • Improving Answer Quality: Ensures that answers are only generated when supported by a complete chain of evidence, reducing hallucinations and unsupported claims.
  • Resource Management: Intelligently manages retrieval calls, preventing unnecessary computations when evidence is already sufficient or when no further retrieval would help.
  • Complex Question Answering: Particularly effective in multi-hop question answering scenarios where evidence sufficiency is critical.