CMSManhattan/JiRackUltra_1b
CMSManhattan's JiRackUltra_1b is a 1.5 billion parameter model, refactored with BitNet features and a redesigned DeepSeek R1 architecture, optimized for efficient CPU inference. It utilizes a unique tokenizer with specialized tags for Routing, Tool call, and Robotics, enabling high-quality tool calling on small models. This model is designed for cost-effective cloud deployments and serves as an expert model in RAG systems, with ready-to-run GGUF quantizations for various performance needs.
Loading preview...
JiRack Ultra 1B: CPU-Optimized Ternary Model
JiRack Ultra 1B is a 1.5 billion parameter model developed by CMSManhattan, specifically engineered for fast and efficient inference on CPUs. It stands out due to its unique architecture and specialized tokenizer, making it highly suitable for specific applications.
Key Capabilities & Features
- BitNet-style Ternary Architecture: Refactored with native BitLinear ternary path (b1.58-style) and $\lambda$-warmup STE, allowing for greater compression and efficient CPU performance.
- Advanced Tokenizer: Features an updated tokenizer with specialized tags for Routing, Tool call, and Robotics, enhancing its ability to perform high-quality tool calling even as a small model.
- CPU Optimization: Designed for cost-effective cloud infrastructure, it offers ready-to-run GGUF quantizations (Q2_K, Q3_K_M, Q4_K_M) to balance size and performance on various CPU setups.
- RAG Expert Model: Positioned as an expert model for Retrieval Augmented Generation (RAG) deployments.
- Custom Quantization Services: Offers custom Quantization-Aware Training (QAT) from user datasets, including double QAT via ONNX, for tailored performance.
Ideal Use Cases
- Tool Calling & Agentic Systems: Excels in scenarios requiring precise tool calls, leveraging its specialized tokenizer for integration with frameworks like ToolBench.
- Cost-Efficient Cloud Deployments: Its CPU optimization and small footprint make it suitable for reducing cloud infrastructure costs.
- Edge & Low-Resource Environments: With various quantizations, it can run effectively on systems with limited RAM (e.g., 2-4 GB).
- Domain-Specific Expert Systems: Can be adapted as a domain-specific tool expert, particularly in robotics and routing applications.