CMSManhattan/JiRackUltra_14b
CMSManhattan/JiRackUltra_14b is a 14.8 billion parameter model built on a DeepSeek R1-14B architecture, refactored with native BitNet ternary features for efficient CPU inference. It utilizes an updated JiRackDeltaNetTokenizer with specialized tags for Routing, Media, Vision, Sound, Tool call, and Robotics, enhancing its capabilities in these domains. Optimized for cloud deployments and RAG systems, this model provides ready-to-run GGUF quantizations for various CPU performance needs. Its unique tokenizer and ternary architecture make it particularly suitable for advanced tool calling and specialized expert roles in resource-constrained environments.
Loading preview...
JiRack Ultra 14B: CPU-Optimized Ternary Model
JiRack Ultra 14B, developed by CMSManhattan, is a 14.8 billion parameter language model engineered for fast and efficient CPU inference. It is built upon a DeepSeek R1-14B architecture and incorporates native BitNet ternary features, allowing for significant compression and optimized performance on CPU hardware. The model comes with ready-to-run GGUF quantizations (Q2_K, Q3_K_M, Q4_K_M) to suit various memory and performance requirements.
Key Capabilities & Differentiators
- Ternary Architecture: Refactored with BitNet features, including a native BitLinear ternary path, enabling high-quality CPU inference with TQ_2 on Llama.cpp and Ollama via Quantization-Aware Training (QAT).
- Advanced JiRackDeltaNetTokenizer: Features an updated tokenizer with specialized tags for Routing, Media, Vision, Sound, Tool call, and Robotics. This enhances its ability to handle complex, domain-specific tasks and advanced tool calling.
- CPU Optimization: Designed for efficient operation on CPUs, making it suitable for cloud-ready deployments and reducing infrastructure costs. It can function as an expert model in RAG deployments.
- Tool Calling Expertise: The specialized tokenizer and architecture are geared towards high-quality tool calling, making it adaptable for agentic or instruct models, particularly as a domain-specific tool expert.
- Flexible Deployment: Supports Docker deployment with various quantization options and is being integrated for production systems on Ollama.
Ideal Use Cases
- Resource-Constrained Environments: Excellent for applications requiring powerful language models on CPU-only infrastructure.
- Advanced Tool Calling & Agents: Its specialized tokenizer makes it highly effective for complex tool integration, robotics control, and routing tasks.
- Domain-Specific Expertise: Well-suited for roles where specialized understanding of media, vision, sound, or robotics is required.
- Cost-Effective Cloud Deployments: Designed to save money on cloud infrastructure due to its CPU optimization and efficient inference.