nehmeailabs-org/nehme-flashcheck-270m

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.3BQuant:BF16Context Size:32kPublished:Dec 15, 2025License:gemmaArchitecture:Transformer0.0K Featherless Exclusive Cold

FlashCheck-270M by Nehme AI Labs is a specialized 0.3 billion parameter Small Language Model (SLM) built on the Gemma 3 270M instruction-tuned family, featuring a 32768 token context length. It is fine-tuned for contextual policy adherence and hallucination detection, acting as a lightweight, privacy-preserving guardrail in RAG pipelines. The model excels at determining if a given claim is fully supported by a document, responding with a definitive 'Yes' or 'No'.

Loading preview...

FlashCheck-270M: Edge Logic Engine

FlashCheck-270M, developed by Nehme AI Labs, is a specialized Small Language Model (SLM) based on the google/gemma-3-270m-it architecture. With 0.3 billion parameters and a 32768 token context length, it is specifically fine-tuned for Contextual Policy Adherence and Hallucination Detection within RAG (Retrieval-Augmented Generation) pipelines.

Key Capabilities

  • Policy Adherence: Determines if a user's claim is consistent with a provided document or policy.
  • Hallucination Detection: Acts as a guardrail to verify factual consistency between a claim and its source document.
  • Binary Output: Provides clear 'Yes' or 'No' verdicts, indicating whether a claim is fully supported by the document.
  • Lightweight & Efficient: Its small size (270M parameters) makes it suitable for edge deployments and privacy-preserving applications.
  • Flexible Deployment: Available in both Transformers format for Python and GGUF for local inference with llama.cpp.

Training & Limitations

FlashCheck was trained on a curated mix of AggreFact-style hallucination detection data and synthetic contrastive policy pairs to enhance its logical reasoning and reduce keyword-matching failures. It is optimized for English language logic and policy checking. The model is intended as a guardrail/verifier and not a general-purpose chat assistant, with output sensitivity to prompt formatting being a known limitation.