CMSManhattan/JiRackUltra_1b

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 2, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

CMSManhattan/JiRackUltra_1b is a 1.5 billion parameter model optimized for efficient CPU inference, featuring a redesigned DeepSeek R1 architecture with native ternary (BitNet-style) support. It incorporates an updated tokenizer with specialized tags for Routing, Tool call, and Robotics, making it particularly effective for agentic applications and domain-specific tool expertise. The model is designed for cloud-ready deployments and offers various GGUF quantizations for balanced performance and memory usage.

Loading preview...

JiRack Ultra 1B: CPU-Optimized Ternary Model

JiRack Ultra 1B, developed by CMSManhattan, is a 1.5 billion parameter model engineered for high-quality CPU inference. It stands out due to its refactored architecture, incorporating BitNet features for native ternary support, which allows for significant compression and efficient operation on standard CPUs.

Key Capabilities & Features

  • Ternary Architecture: Built on a redesigned DeepSeek R1 architecture with native ternary (BitNet-style) support, enabling high-quality CPU inference, especially with TQ_2 quantizations via Llama.cpp and Ollama.
  • Advanced Tokenizer: Features an updated tokenizer with specialized tags for Routing, Tool call, and Robotics, enhancing its capabilities for agentic workflows and domain-specific tasks.
  • CPU Optimization: Designed for fast and efficient inference on CPUs, making it suitable for cloud-ready deployments and cost-effective infrastructure.
  • GGUF Quantizations: Available with ready-to-run GGUF quantizations (Q2_K, Q3_K_M, Q4_K_M) to balance performance and memory footprint, with options for further ternary compression via custom QAT.
  • Tool Calling Expertise: Optimized for tool calling, leveraging its unique tokenizer to enable high-quality function calling in small models, acting as a domain-specific tool expert.

Ideal Use Cases

  • Agentic Applications: Excellent for developing coding agents, robotics control, and routing systems due to its specialized tokenizer and tool-calling capabilities.
  • Cost-Effective Deployments: Its CPU optimization and efficient architecture make it a strong candidate for cloud deployments where cost savings on infrastructure are critical.
  • Edge & Local Inference: Suitable for running on consumer-grade hardware, including laptops and SBCs, with various quantization options to match memory constraints.
  • RAG Deployments: Can serve as an expert model in Retrieval Augmented Generation (RAG) systems, particularly when integrated with the ONNX JiRack Java server.