startlux-models/StartLux-Decision-2B

VISIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2.3BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 29, 2026License:cc-by-nc-4.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

StartLux-Decision-2B is a 2.3 billion parameter model developed by StartLux Labs, designed to answer typed questions about a given state by providing a probability for each option. It supports various input types including text, JSON, and images, with a substantial context length of up to 262,144 tokens. This model specializes in decision-making tasks, offering probabilistic outputs for choice, yes/no, and rating questions, and is optimized for fast inference with end-to-end latencies as low as 9.6 ms for a single yes/no question.

Loading preview...

StartLux-Decision-2B Overview

StartLux-Decision-2B, developed by StartLux Labs, is a 2.3 billion parameter model specifically engineered for probabilistic decision-making. It processes typed questions about a given state (text, JSON, or images) and returns a probability for each possible option, supporting choice, yes/no, and rating question types. A key differentiator is its ability to handle extremely long inputs, with a native context length of up to 262,144 tokens, allowing for comprehensive analysis of complex states.

Key Capabilities

  • Probabilistic Decision Outputs: Provides a probability score for each option in response to questions, enabling nuanced decision support.
  • Multi-modal Input: Accepts text, JSON, and image inputs, with a built-in vision tower for image processing.
  • Extended Context Window: Supports an impressive 262,144-token context length, facilitating the processing of very long documents or complex data structures.
  • High-Speed Inference: Optimized for fast inference, achieving end-to-end latencies of 15.5 ms for three questions and 9.6 ms for a single yes/no question on an H200 GPU, utilizing fast kernels and CUDA graph replays.
  • TypeSafe API Compatibility: Uses the TypeSafe /v1/systemone format, ensuring compatibility with existing clients.

Good For

  • Applications requiring probabilistic answers to structured questions.
  • Analyzing large volumes of text, JSON, or image data for decision support.
  • Use cases where low-latency decision inference is critical.
  • Integrating with systems already using the TypeSafe /v1/systemone API.