CMSManhattan/JiRackUltra_1b
JiRack Ultra 1B is a 1.5 billion parameter model from CMSManhattan, optimized for efficient CPU inference. Built on a redesigned DeepSeek R1 architecture with native ternary (BitNet-style) support, it features an updated tokenizer with specialized tags for Routing, Tool calls, and Robotics. This model excels as a domain-specific expert for agentic and instruct tasks, particularly in tool calling and robotics applications, offering high-quality CPU inference with various GGUF quantizations.
Loading preview...
JiRack Ultra 1B: CPU-Optimized Ternary Model for Agentic AI
JiRack Ultra 1B is a 1.5 billion parameter language model developed by CMSManhattan, specifically engineered for fast and efficient CPU inference. It leverages a redesigned DeepSeek R1 architecture with native ternary (BitNet-style) support, enabling significant compression and high-quality performance on CPU-bound systems.
Key Capabilities
- CPU-Optimized Performance: Designed for efficient inference on CPUs, offering excellent interactive speeds even on modest hardware (e.g., Ryzen 5 / Intel i5 with 4-8 GB RAM).
- Ternary Architecture: Incorporates BitNet features for advanced compression, with ready-to-run GGUF quantizations (Q2_K, Q3_K_M, Q4_K_M) and options for further QAT (Quantization-Aware Training) for specific tasks.
- Specialized Tokenizer: Features an updated tokenizer with unique tags for Routing, Tool calls, and Robotics, enhancing its capabilities for agentic workflows and domain-specific applications.
- Tool Calling Expertise: Excels at high-quality tool calling, making it suitable for integrating with external functions and APIs, including support for Spring Boot AI and GoEx AI tool call examples.
- Robotics Integration: Advanced tokenizer features specifically designed for robotics applications, supporting the ROBOS HUMANOID Project.
Good For
- Cost-Effective AI Deployments: Ideal for cloud-ready deployments where minimizing infrastructure costs is crucial, as it performs well on standard CPU resources.
- Agentic and Instruct Models: Particularly effective as a domain-specific expert for agentic or instruct models that require precise tool calling and routing capabilities.
- Edge and Embedded Systems: Its small size and CPU optimization make it a strong candidate for deployment on devices with limited resources, such as laptops or SBCs.
- Robotics and Automation: Suited for applications requiring advanced control and interaction through its specialized robotics and routing tags.
- Developers Seeking Custom Quantization: Offers flexibility for custom QAT to achieve optimal balance between model size and performance for specific datasets and tasks.