kirp/jpt-9b

VISIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 25, 2026License:cc-by-nc-4.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

kirp/jpt-9b is a 9 billion parameter open decision model, fine-tuned by kirp on Qwen/Qwen3.5-9B with a 32768 token context length. It is specifically designed to provide calibrated probabilities for typed questions (choice, score, noul) given a situation, focusing on fast, single-pass inference without generating explanations or reasoning tokens. This model excels at general typed decision-making, achieving a Decision Index 0.2.1 score of 46.89, making it the best 9B model on that benchmark.

Loading preview...

JPT-9B: A Fast, Open Decision Model

JPT-9B, developed by kirp, is a 9 billion parameter model built on Qwen/Qwen3.5-9B, designed for rapid, open decision-making. Unlike traditional LLMs, it provides calibrated probabilities for typed questions (choice, score, noul) in a single forward pass, without generating explanatory text. This architecture prioritizes low latency, making it suitable for applications requiring immediate probabilistic answers.

Key Capabilities

  • Typed Decision Interface: Implements the typed-decision interface, accepting choice (2-255 labels), score (ordered scale), and noul (yes/no) questions.
  • Fast Inference: Optimized for speed, delivering probabilities with one prefill latency, as it avoids generating reasoning tokens.
  • Strong Decision Benchmarks: Achieves a JevBench v1.4.0 public accuracy of 0.853 and a Decision Index 0.2.1 score of 46.89, positioning it as the top 9B model on the Decision Index.
  • Vision-Capable: Built on Qwen3.5-9B, its vision tower remains unchanged, allowing for potential image-based decision tasks (though untested on this specific checkpoint).

Good For

  • Applications requiring probabilistic answers: Ideal for scenarios where a calibrated probability for a specific decision is needed, rather than a generated explanation.
  • Low-latency decision systems: Its single-pass inference makes it suitable for real-time or high-throughput decision engines.
  • Integration with llm2jev: Designed to work seamlessly with the llm2jev tool for serving and querying, supporting backends like SGLang and vLLM.
  • Non-commercial research and development: Licensed under CC BY-NC 4.0, it's available for non-commercial use in exploring advanced decision-making AI.