h2oai/h2o-lightning-4b

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Oct 2, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

H2O-Lightning-4B is a 4.5 billion parameter decision model developed by H2O.ai, built upon Qwen/Qwen3.5-4B, with a 32K context length. It specializes in answering typed decision questions about records and associated images, returning probabilities for choices, yes/no statements, and ordinal scores. This model achieved the #1 composite score on JevBench v1.6.1, outperforming larger models in decision-making efficiency and calibration.

Loading preview...

H2O-Lightning-4B: A Specialized Decision Model

H2O-Lightning-4B, developed by H2O.ai, is a 4.5 billion parameter model based on Qwen/Qwen3.5-4B, designed for efficient and calibrated decision-making. It processes typed questions about records, documents, tickets, and images, providing probabilistic answers for choices, yes/no statements, and ordinal scores. The model runs on unmodified vLLM 0.30.0, with each decision requiring only one forward pass and one output token, optimizing for cost and latency.

Key Capabilities

  • Top Performance on JevBench: Achieved the #1 composite score (72.5) on JevBench v1.6.1, surpassing Jev itself (71.5) and larger 12B, 26B, and 31B open models, particularly excelling in calibration and speed.
  • Multimodal Decision-Making: Capable of processing both text and image inputs to answer decision questions, supporting up to 4 images per request, each scaled to 1.6 megapixels.
  • High Calibration: Designed for probabilities that accurately reflect confidence, a first-class goal in its development.
  • Efficient Deployment: Runs with a small standard-library shim, enabling hybrid mode serving alongside the base Qwen3.5-4B model for both decisions and free-text generation from a single vLLM instance.

Good for

  • Automated decision-making in applications like processing claims, categorizing tickets, or evaluating policies.
  • Use cases requiring highly calibrated probabilistic outputs for classification and scoring tasks.
  • Integrating multimodal (text and image) analysis into decision workflows.
  • Environments where low latency and cost-effective inference for decision tasks are critical.