CMSManhattan/JiRackUltra_32b
CMSManhattan/JiRackUltra_32b is a 32.8 billion parameter model built on a DeepSeek R1-32B architecture, refactored with native BitNet ternary features for efficient CPU inference. It utilizes an updated tokenizer with specialized tags for Routing, Media, Vision, Sound, Tool call, and Robotics, enhancing its capabilities in these domains. Optimized for cloud deployments, JiRackUltra_32b excels as an expert model in RAG systems and advanced tool-calling scenarios, offering ready-to-run GGUF quantizations for various CPU environments.
Loading preview...
JiRack Ultra 32B: CPU-Optimized Ternary Model
JiRack Ultra 32B, developed by CMSManhattan, is a 32.8 billion parameter model engineered for high-quality, efficient inference on CPUs. It is built upon a DeepSeek R1-32B architecture and incorporates native BitNet ternary features, allowing for significant compression and performance benefits, particularly with GGUF quantizations like Q2_K, Q3_K_M, and Q4_K_M.
Key Differentiators & Capabilities
- Ternary Architecture: Refactored with BitNet features (native BitLinear ternary path) for optimized CPU inference and further compression potential via Quantization-Aware Training (QAT).
- Advanced Tokenizer: Features the JiRackDeltaNetTokenizer, which includes specialized tags for Routing, Media, Vision, Sound, Tool call, and Robotics. This enhances its ability to handle complex, domain-specific tasks.
- Cloud & RAG Optimization: Designed to be cloud-ready, reducing infrastructure costs. It is particularly effective as an expert model within RAG (Retrieval-Augmented Generation) deployments.
- Enhanced Tool Calling: The specialized tokenizer and architecture enable high-quality tool calling, making it suitable for agentic workflows and integration with tool libraries like Spring Boot AI and GoEx AI.
- CPU Performance: Provides ready-to-run GGUF quantizations optimized for various CPU configurations, with support for AVX2 and AVX-512 instructions for high performance.
Ideal Use Cases
- CPU-constrained environments: When GPU resources are limited or cost-prohibitive.
- RAG deployments: As a specialized expert model for specific domains.
- Advanced tool calling and agentic systems: Leveraging its unique tokenizer for precise function calling in robotics, routing, and multimedia applications.
- Edge deployments: Where efficient, compact models are required.