saidutta69/clef-flash-heretic

VISIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 4, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The saidutta69/clef-flash-heretic is a 9B parameter multimodal decision model, built on a Qwen3.5 backbone, designed to classify and score inputs without generating free text. This variant, produced using Heretic v1.4.0, significantly reduces refusal behavior from 99/100 to 63/100 on harmful evaluation sets through targeted weight edits, rather than fine-tuning. It excels in classification and decision pipelines for tasks like moderation triage, risk scoring, and content classification, especially when dealing with inputs the original model would refuse, and supports a 32768 token context length.

Loading preview...

Overview

The saidutta69/clef-flash-heretic is a 9-billion parameter multimodal decision model based on the Qwen3.5 architecture, specifically a decensored variant of Cloudflare/clef-flash. It was created using Heretic v1.4.0, which employs directional ablation (abliteration) to suppress refusal behavior through targeted weight edits to the attention output and MLP down-projections, rather than traditional fine-tuning. This approach maintains the original model's Decision Index scoring and typed-output head while reducing refusals.

Key Capabilities

  • Refusal Suppression: Reduces refusals on harmful evaluation sets from 99/100 to 63/100, allowing it to score cases the original model would decline.
  • Multimodal Decision Making: Reads state (text, JSON, image, or video) and a schema of typed questions, returning probabilities for every allowed option in a single forward pass.
  • No Free-Form Text Generation: Designed specifically for classification and decision tasks, it produces per-option probabilities without generating free text, eliminating the need for output parsing.
  • High Fidelity Edits: Achieves a KL divergence of 0.0189, indicating minimal collateral damage to the model's core option-scoring behavior while removing refusal directions.
  • Local Deployment: Can run locally via GGUF on GPUs with as little as 16 GB, with options for quantization to fit smaller GPUs or CPU-only setups.

Good For

  • Engineers building classification and decision pipelines who encounter refusals at inference time.
  • Applications requiring moderation triage, routing, risk scoring, and content classification.
  • Use cases where a model needs to score inputs that might typically trigger refusal behavior in standard models.
  • Integration with existing systems like Jev and SystemOne, with which its API is compatible.