ulamai/Ulam-1

VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 16, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Ulam-1 is a 27.357 billion parameter mathematical reasoning model developed by Ulam AI, built upon the Qwen3.8-27B architecture. This model is specifically post-trained for exploratory mathematical problem solving, proof debugging, and research assistance. It excels in generating detailed mathematical reasoning and is intended for applications requiring robust mathematical inference with human oversight. Ulam-1 supports a validated serving context of up to 131,072 tokens, making it suitable for complex mathematical problems.

Loading preview...

Ulam-1: A Specialized Mathematical Reasoning Model

Ulam-1 is a 27.357 billion parameter model developed by Ulam AI, derived from the Qwen3.8-27B base model through cumulative post-training. It is distributed as a standalone merged BF16 Transformers model, requiring no separate Qwen checkpoint or adapter. The model's post-training focuses on modifying low-rank modules in the text decoder while preserving the upstream tokenizer, processor, chat template, and vision branch.

Key Capabilities

  • Advanced Mathematical Reasoning: Specifically designed for exploratory mathematical problem solving, proof debugging, and statement checking.
  • Research Assistance: Aids in summarizing research progress and suggesting candidate approaches.
  • Multimodal Architecture: Retains the upstream Qwen multimodal architecture, supporting text-only, image, and video messages, though Ulam post-training is text-focused.
  • Extended Context Length: Supports a validated serving context of 131,072 tokens, with an architectural setting of 262,144 tokens, suitable for long mathematical generations.
  • Explicit Thinking Spans: Operates in a "thinking mode" by default, capable of emitting explicit <think>...</think> spans for detailed reasoning.

Evaluation and Training

Ulam-1's development involved a multi-stage training lineage, including supervised research-reasoning refinement, verified-reward optimization, and conclusion-focused mathematical reinforcement learning. Evaluation on ErdosBench, a correctness-first signal for research mathematics, showed strong performance, with 137 out of 226 problems receiving A or B grades in a solve-only audit.

Intended Uses

  • Exploratory mathematical problem solving with human review.
  • Proof debugging, obstruction finding, and statement checking.
  • Research-progress summaries and candidate approaches.
  • Research on verifier-backed mathematical post-training.
  • Text-focused inference and experimental multimodal mathematical assistance.

Limitations

It is crucial to note that Ulam-1's outputs are not formal proofs and may contain mathematical errors, requiring independent expert review. The model's evaluation is based on a single deterministic generation and item-level judgment pass, and it does not provide sampling-variance estimates or inter-rater agreement statistics.