Cloudflare/clef-flash
Cloudflare's Clef-Flash is a 9 billion parameter multimodal model, post-trained from Qwen/Qwen3.5-9B, designed for decision-making tasks. It processes state information from text, JSON, images, or video inputs and outputs probabilities for predefined options across multiple questions in a single forward pass. This model specializes in structured decision-making without free-form text generation, making it ideal for applications requiring precise, probabilistic answers to specific queries.
Loading preview...
Overview
Clef-Flash is a 9 billion parameter multimodal model developed by Cloudflare, built upon the Qwen/Qwen3.5-9B backbone. Unlike traditional LLMs that generate free-form text, Clef-Flash is specifically engineered to convert a given state (which can include text, JSON, images, or video) and a schema of typed questions into concrete decisions. It provides a probability for every allowed option of every question in a single pass, eliminating the need for output parsing.
Key Capabilities
- Multimodal Input Processing: Accepts diverse input types including text, JSON, images, and video to inform decision-making.
- Structured Decision Output: Generates probabilities for predefined options across multiple questions, rather than free-form text.
- Efficient Inference: Processes all questions in a single forward pass, designed for speed and direct probabilistic answers.
- Jev/SystemOne API Compatibility: Fully compatible with Jev and SystemOne APIs for seamless integration into existing workflows.
- Specialized Architecture: Utilizes a Qwen/Qwen3.5-9B backbone with a small transformer head for joint schema processing, routing evidence, and scoring options.
Performance Highlights
Clef-Flash demonstrates strong performance across various benchmarks, particularly excelling in specific decision-making tasks. On the Decision Index leaderboard, it shows competitive or leading results in categories such as BFCL (98.8%), API-Bank (93.1%), ContractNLI (84.3%), and ForecastBench (10.6 Brier score, lower is better). It also achieves high accuracy in several ARC and MMLU benchmarks. Notably, Clef-Flash boasts a significantly lower median latency (38.8 ms) compared to many other models in its class, making it suitable for real-time applications.
When to Use Clef-Flash
This model is particularly well-suited for use cases requiring precise, probabilistic answers to structured questions based on diverse inputs. It's an excellent choice for applications like:
- Automated invoice processing and status determination.
- Customer service routing based on incident descriptions.
- Content moderation or classification where specific criteria need to be evaluated.
- Any scenario where a system needs to make a clear, quantifiable decision from multimodal data without generating open-ended responses.