CMSManhattan/JiRackUltra_1b

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 2, 2026License:mitArchitecture:Transformer0.0K Open Weights Featherless Exclusive Cold

JiRack Ultra 1B is a 1.5 billion parameter model from CMSManhattan, optimized for efficient CPU inference. Built on a redesigned DeepSeek R1 architecture with native ternary (BitNet-style) support, it features an updated tokenizer with specialized tags for Routing, Tool calls, and Robotics. This model excels as a domain-specific expert for agentic and instruct tasks, particularly in tool calling and robotics applications, offering high-quality CPU inference with various GGUF quantizations.

Loading preview...

JiRack Ultra 1B: CPU-Optimized Ternary Model for Agentic AI

JiRack Ultra 1B is a 1.5 billion parameter language model developed by CMSManhattan, specifically engineered for fast and efficient CPU inference. It leverages a redesigned DeepSeek R1 architecture with native ternary (BitNet-style) support, enabling significant compression and high-quality performance on CPU-bound systems.

Key Capabilities

  • CPU-Optimized Performance: Designed for efficient inference on CPUs, offering excellent interactive speeds even on modest hardware (e.g., Ryzen 5 / Intel i5 with 4-8 GB RAM).
  • Ternary Architecture: Incorporates BitNet features for advanced compression, with ready-to-run GGUF quantizations (Q2_K, Q3_K_M, Q4_K_M) and options for further QAT (Quantization-Aware Training) for specific tasks.
  • Specialized Tokenizer: Features an updated tokenizer with unique tags for Routing, Tool calls, and Robotics, enhancing its capabilities for agentic workflows and domain-specific applications.
  • Tool Calling Expertise: Excels at high-quality tool calling, making it suitable for integrating with external functions and APIs, including support for Spring Boot AI and GoEx AI tool call examples.
  • Robotics Integration: Advanced tokenizer features specifically designed for robotics applications, supporting the ROBOS HUMANOID Project.

Good For

  • Cost-Effective AI Deployments: Ideal for cloud-ready deployments where minimizing infrastructure costs is crucial, as it performs well on standard CPU resources.
  • Agentic and Instruct Models: Particularly effective as a domain-specific expert for agentic or instruct models that require precise tool calling and routing capabilities.
  • Edge and Embedded Systems: Its small size and CPU optimization make it a strong candidate for deployment on devices with limited resources, such as laptops or SBCs.
  • Robotics and Automation: Suited for applications requiring advanced control and interaction through its specialized robotics and routing tags.
  • Developers Seeking Custom Quantization: Offers flexibility for custom QAT to achieve optimal balance between model size and performance for specific datasets and tasks.