CMSManhattan/JiRackUltra_7b
JiRack Ultra 7B by CMSManhattan is a 7.6 billion parameter model built on a DeepSeek R1-7B architecture, refactored with BitNet features for native ternary support and optimized for CPU inference. It utilizes an updated tokenizer with specialized tags for Routing, Media, Vision, Sound, Tool call, and Robotics, making it highly efficient for domain-specific expert applications and cloud infrastructure cost savings. The model provides ready-to-run GGUF quantizations and is designed for high-quality CPU inference, particularly in agentic and instruct models for advanced tool calling.
Loading preview...
JiRack Ultra 7B: CPU-Optimized Ternary Model
JiRack Ultra 7B is a 7.6 billion parameter model developed by CMSManhattan, built upon a DeepSeek R1-7B architecture. Its core differentiator is the integration of BitNet features for native ternary (BitNet-style) support, enabling highly efficient CPU inference. The model is designed to be cloud-ready, aiming to reduce infrastructure costs, and is particularly effective as an expert model in RAG deployments.
Key Capabilities & Features
- Ternary Architecture: Refactored with BitNet features for high-quality CPU inference, supporting TQ_2 on Llama.cpp and Ollama via QAT.
- Advanced Tokenizer: Features an updated tokenizer with specialized tags for Routing, Media, Vision, Sound, Tool call, and Robotics, enhancing its capabilities in these domains.
- CPU Optimization: Provides ready-to-run GGUF quantizations (Q2_K, Q3_K_M, Q4_K_M) for various memory and performance needs, with support for AVX2 and AVX-512 CPU instructions.
- Tool Calling Expertise: Adapts to agentic or instruct models for high-quality tool calling, leveraging its unique tokenizer for domain-specific tool expertise.
- Commercial Support: Offers services for custom ternary compression (QAT), adaptation to avoid catastrophic forgetting, and deployment/scaling solutions.
Ideal Use Cases
- Cost-Efficient Cloud Deployments: Optimized for CPU inference to save on cloud infrastructure expenses.
- Domain-Specific Expert Systems: Excels in applications requiring specialized understanding of Routing, Media, Vision, Sound, Tool calling, and Robotics.
- Agentic & Instruct Models: Particularly suited for scenarios demanding advanced and high-quality tool calling capabilities.
- Edge & Low-Resource Environments: Its efficient quantizations make it viable for deployment on systems with limited RAM, such as laptops or edge devices.