CMSManhattan/JiRackUltra_7b
CMSManhattan/JiRackUltra_7b is a 7.6 billion parameter model built on a DeepSeek R1 -7B architecture, refactored with native BitNet ternary features for efficient CPU inference. It utilizes an updated tokenizer with specialized tags for Routing, Media, Vision, Sound, Tool call, and Robotics, making it highly optimized for agentic workflows and domain-specific tool calling. This model is designed for high-quality CPU inference and cost-effective cloud deployments, particularly as an expert model in RAG systems.
Loading preview...
JiRack Ultra 7B: CPU-Optimized Ternary Model for Agentic AI
JiRack Ultra 7B is a 7.6 billion parameter model developed by CMSManhattan, engineered for fast and efficient CPU inference. It is built upon a DeepSeek R1 -7B architecture, incorporating native ternary (BitNet-style) features and an updated tokenizer. This unique architecture enables high-quality CPU inference, particularly with TQ_2 quantizations on Llama.cpp and Ollama via Quantization-Aware Training (QAT).
Key Capabilities & Features
- Ternary Architecture: Refactored with BitNet features for enhanced CPU efficiency and potential for further compression.
- Advanced Tokenizer: Features the JiRackDeltaNetTokenizer with specialized tags for Routing, Media, Vision, Sound, Tool call, and Robotics, enabling high-quality tool calling even in smaller models.
- CPU Optimization: Designed for cost-effective cloud infrastructure and efficient local deployment, with ready-to-run GGUF quantizations (Q2_K, Q3_K_M, Q4_K_M).
- Agentic AI Focus: Excels in agentic or instruct models for tool calling, functioning as a domain-specific tool expert.
- Cloud-Ready: Positioned as a cloud-ready model, suitable for RAG deployments and offering an ONNX JiRack Java server alternative.
Good for
- CPU-constrained environments: Ideal for deployments where GPU resources are limited or costly.
- Agentic workflows: Particularly strong in applications requiring advanced tool calling, routing, and robotics integration.
- Cost-sensitive cloud deployments: Helps reduce infrastructure costs due to its CPU optimization.
- Domain-specific expert systems: Can be tailored via QAT for specific tasks and datasets to avoid catastrophic forgetting.
- Edge computing: Usable on lower-spec hardware like laptops with acceptable performance.