juspay/Xor-26B-A4B

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:2Model Size:26BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 1, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

juspay/Xor-26B-A4B is a 26 billion parameter Mixture-of-Experts (MoE) model, post-trained by Juspay from Google's Gemma-4-26B-A4B-it, specifically for typed decision tasks. It features 30 layers with 128 routed experts (8 active per token) and is delivered as merged BF16 weights. This model excels at answering 'noul', 'choice', and 'score' questions with full probability distributions, making it suitable for applications requiring precise, probabilistic decision-making.

Loading preview...

Overview

juspay/Xor-26B-A4B is a 26 billion parameter Mixture-of-Experts (MoE) model, post-trained by Juspay from Google's Gemma-4-26B-A4B-it. It is specifically designed for typed decision tasks, providing full probability distributions for noul (binary probability), choice (categorical decision), and score (expected ordinal score) questions. The model is built on a base with 30 layers and 128 routed experts, with approximately 4 billion active parameters per token, and is released as merged BF16 weights.

Key Capabilities

  • Specialized for Decision Tasks: Optimized to answer noul, choice, and score questions, providing detailed probability distributions for each.
  • Mixture-of-Experts Architecture: Leverages a Gemma-4-26B-A4B-it base with 128 experts, enhancing efficiency and performance for its target tasks.
  • TypeSafe-compatible API: Served through a /v1/systemone API, ensuring structured and reliable interaction for decision-making applications.
  • Reproducible Results: Designed for consistent output, with a bundled compatibility server and specific temperature settings for choice (1.95) and noul/score (1.0).
  • BF16 Weights: Provided in BF16 format, with a 4-bit quantized version (Xor-26B-A4B-NVFP4) also available, offering a balance between performance and memory footprint.

Good for

  • Automated Decision Systems: Ideal for applications requiring probabilistic answers to structured questions.
  • Categorical and Ordinal Scoring: Excels in scenarios where a model needs to provide not just a choice, but also the confidence or distribution across options.
  • Research and Development: Useful for exploring MoE architectures and post-training techniques on a Gemma base for specific task optimization.
  • Applications requiring high reproducibility: The model's design and serving setup emphasize consistent and verifiable outputs for critical decision tasks.