startlux-models/StartLux-Decision-2B
StartLux-Decision-2B is a 2.3 billion parameter model developed by StartLux Labs, designed to answer typed questions about a given state by providing a probability for each option. It supports various input types including text, JSON, and images, with a substantial context length of up to 262,144 tokens. This model specializes in decision-making tasks, offering probabilistic outputs for choice, yes/no, and rating questions, and is optimized for fast inference with end-to-end latencies as low as 9.6 ms for a single yes/no question.
Loading preview...
StartLux-Decision-2B Overview
StartLux-Decision-2B, developed by StartLux Labs, is a 2.3 billion parameter model specifically engineered for probabilistic decision-making. It processes typed questions about a given state (text, JSON, or images) and returns a probability for each possible option, supporting choice, yes/no, and rating question types. A key differentiator is its ability to handle extremely long inputs, with a native context length of up to 262,144 tokens, allowing for comprehensive analysis of complex states.
Key Capabilities
- Probabilistic Decision Outputs: Provides a probability score for each option in response to questions, enabling nuanced decision support.
- Multi-modal Input: Accepts text, JSON, and image inputs, with a built-in vision tower for image processing.
- Extended Context Window: Supports an impressive 262,144-token context length, facilitating the processing of very long documents or complex data structures.
- High-Speed Inference: Optimized for fast inference, achieving end-to-end latencies of 15.5 ms for three questions and 9.6 ms for a single yes/no question on an H200 GPU, utilizing fast kernels and CUDA graph replays.
- TypeSafe API Compatibility: Uses the TypeSafe
/v1/systemoneformat, ensuring compatibility with existing clients.
Good For
- Applications requiring probabilistic answers to structured questions.
- Analyzing large volumes of text, JSON, or image data for decision support.
- Use cases where low-latency decision inference is critical.
- Integrating with systems already using the TypeSafe
/v1/systemoneAPI.