CMSManhattan/JiRackUltra_14b
CMSManhattan/JiRackUltra_14b is a 14.8 billion parameter language model built on a DeepSeek R1-14B architecture, optimized for CPU inference and featuring native ternary (BitNet-style) support. It incorporates an updated tokenizer with specialized tags for Routing, Media, Vision, Sound, Tool call, and Robotics, making it highly suitable for agentic applications, advanced tool calling, and robotics. The model is designed for efficiency and cost-effectiveness in cloud deployments, offering ready-to-run GGUF quantizations for various hardware configurations.
Loading preview...
JiRack Ultra 14B: CPU-Optimized Ternary Model for Agentic AI
CMSManhattan/JiRackUltra_14b is a 14.8 billion parameter model engineered for efficient CPU inference, leveraging a DeepSeek R1-14B architecture with native ternary (BitNet-style) support. This model stands out due to its specialized tokenizer, which includes unique tags for Routing, Media, Vision, Sound, Tool call, and Robotics, significantly enhancing its capabilities in domain-specific applications.
Key Capabilities
- CPU-Optimized Performance: Designed for high-quality inference on CPUs, making it cost-effective for cloud infrastructure and local deployments.
- Advanced Tool Calling: The updated JiRack Precision Tokenizer enables sophisticated tool calling, supporting integration with frameworks like ToolBench and Spring Boot AI.
- Robotics and Agentic AI: Specialized tags and architecture make it highly suitable for robotics control, routing, and complex agentic workflows.
- Ternary Architecture: Incorporates BitNet features for efficient quantization (TQ_2 on Llama.cpp and Ollama), allowing for significant compression and performance benefits.
- Flexible Quantizations: Available in various GGUF quantizations (Q4_K_M, Q3_K_M, Q2_K) to balance performance and memory footprint across different hardware.
Good For
- Cost-Effective Cloud Deployments: Its CPU optimization and efficient architecture reduce infrastructure costs.
- Agentic Applications: Ideal for building intelligent agents requiring advanced tool use, routing, and decision-making.
- Robotics and Automation: Suited for tasks involving robotic control and complex automated systems due to its specialized tokenizer.
- Local Development and Edge Devices: GGUF quantizations and CPU focus make it accessible for local development, even on systems with limited GPU resources.
- Domain-Specific Expert Models: Can serve as an expert model in RAG deployments, particularly where specialized understanding of routing, media, or robotics is required.