startlux-models/StartLux-Decision-9B

VISIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 29, 2026License:cc-by-nc-4.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

StartLux-Decision-9B is a 9 billion parameter model from startlux-models designed for structured decision-making tasks. It processes typed questions (choice, yes/no, rating) about a given state, which can be text, JSON, or images, and returns probabilities for each option. The model supports an extensive context length of up to 262,144 tokens and is optimized for fast inference, achieving low latencies on H200 GPUs, making it suitable for high-throughput decision-making applications.

Loading preview...

StartLux-Decision-9B Overview

StartLux-Decision-9B is a 9 billion parameter model developed by startlux-models, specifically engineered for structured decision-making. Unlike general-purpose LLMs, it excels at answering typed questions (e.g., multiple choice, yes/no, rating scales) about a given state, providing a probability for each potential option. This model is part of a family of decision models, offering various sizes to suit different performance and latency requirements.

Key Capabilities

  • Typed Question Answering: Processes questions about a state and returns probabilistic answers for predefined options.
  • Multimodal Input: Accepts state descriptions in text, JSON, or images, with a built-in vision tower for image processing.
  • Extended Context Length: Supports exceptionally long inputs up to 262,144 tokens (256K), allowing for comprehensive context analysis.
  • High-Speed Inference: Optimized with custom kernels (flash-linear-attention, causal-conv1d) for fast, end-to-end inference, achieving latencies as low as 17.6 ms for a single yes/no question on an H200 GPU.
  • Batch Processing: Features decide_batch for efficient processing of multiple requests simultaneously, significantly improving throughput.
  • TypeSafe API Compatibility: Uses the /v1/systemone format, ensuring compatibility with existing Jev clients.

Good For

  • Automated Triage Systems: Categorizing customer support tickets, assigning severity, or routing to appropriate teams based on textual or image-based evidence.
  • Probabilistic Decision Support: Applications requiring a confidence score or probability distribution for various choices, rather than a single definitive answer.
  • High-Throughput Decision Engines: Scenarios where rapid, low-latency decision-making is critical, such as real-time analytics or operational automation.
  • Complex Data Analysis: Analyzing extensive textual or multimodal inputs to derive structured decisions, leveraging its large context window.