Cloudflare/clef-flash

Hugging Face
VISIONPricing:Input $0.1078 / Cached $0.0862 / Output $0.28Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 30, 2026License:apache-2.0Architecture:Transformer0.4K Open Weights Warm

Cloudflare's Clef-Flash is a 9 billion parameter multimodal model, post-trained from Qwen/Qwen3.5-9B, designed for decision-making tasks. It processes state information from text, JSON, images, or video inputs and outputs probabilities for predefined options across multiple questions in a single forward pass. This model specializes in structured decision-making without free-form text generation, making it ideal for applications requiring precise, probabilistic answers to specific queries.

Loading preview...

Overview

Clef-Flash is a 9 billion parameter multimodal model developed by Cloudflare, built upon the Qwen/Qwen3.5-9B backbone. Unlike traditional LLMs that generate free-form text, Clef-Flash is specifically engineered to convert a given state (which can include text, JSON, images, or video) and a schema of typed questions into concrete decisions. It provides a probability for every allowed option of every question in a single pass, eliminating the need for output parsing.

Key Capabilities

  • Multimodal Input Processing: Accepts diverse input types including text, JSON, images, and video to inform decision-making.
  • Structured Decision Output: Generates probabilities for predefined options across multiple questions, rather than free-form text.
  • Efficient Inference: Processes all questions in a single forward pass, designed for speed and direct probabilistic answers.
  • Jev/SystemOne API Compatibility: Fully compatible with Jev and SystemOne APIs for seamless integration into existing workflows.
  • Specialized Architecture: Utilizes a Qwen/Qwen3.5-9B backbone with a small transformer head for joint schema processing, routing evidence, and scoring options.

Performance Highlights

Clef-Flash demonstrates strong performance across various benchmarks, particularly excelling in specific decision-making tasks. On the Decision Index leaderboard, it shows competitive or leading results in categories such as BFCL (98.8%), API-Bank (93.1%), ContractNLI (84.3%), and ForecastBench (10.6 Brier score, lower is better). It also achieves high accuracy in several ARC and MMLU benchmarks. Notably, Clef-Flash boasts a significantly lower median latency (38.8 ms) compared to many other models in its class, making it suitable for real-time applications.

When to Use Clef-Flash

This model is particularly well-suited for use cases requiring precise, probabilistic answers to structured questions based on diverse inputs. It's an excellent choice for applications like:

  • Automated invoice processing and status determination.
  • Customer service routing based on incident descriptions.
  • Content moderation or classification where specific criteria need to be evaluated.
  • Any scenario where a system needs to make a clear, quantifiable decision from multimodal data without generating open-ended responses.