CMSManhattan/JiRackUltra_14b

TEXT GENERATIONConcurrent Unit Cost:1Model Size:14.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 3, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

CMSManhattan/JiRackUltra_14b is a 14 billion parameter model based on a DeepSeek R1-14B architecture, refactored with native BitNet features for efficient CPU inference. It includes an updated tokenizer with specialized tags for Routing, Media, Vision, Sound, Tool call, and Robotics. Optimized for cloud-ready deployments and cost savings, this model is available with ready-to-run GGUF quantizations for various memory footprints.

Loading preview...

JiRack Ultra 14B Overview

CMSManhattan/JiRackUltra_14b is a 14 billion parameter language model built on a DeepSeek R1-14B architecture, specifically engineered for fast and efficient CPU inference. A key differentiator is its integration of BitNet features, providing a native ternary (b1.58-style) path with λ-warmup STE, which allows for greater compression than standard FP16 models. The model's tokenizer has been significantly updated to include new specialized tags for Routing, Media, Vision, Sound, Tool call, and Robotics, enhancing its capabilities for diverse, multimodal applications.

Key Capabilities & Features

  • CPU Optimization: Designed for efficient performance on standard CPUs, making it suitable for cloud-ready deployments and reducing infrastructure costs.
  • Ternary Architecture: Refactored with BitNet features, enabling advanced compression and the potential for custom Quantization-Aware Training (QAT) for specific datasets.
  • Extended Tokenizer: Supports a wide range of specialized tags, indicating potential for advanced routing, multimodal processing, and tool interaction.
  • GGUF Quantizations: Available in various GGUF quantizations (Q2_K, Q3_K_M, Q4_K_M) to balance model size, RAM usage, and performance, with options from 6.81 GB (Q2_K) to 28.1 GB (Full precision).
  • Docker Integration: Provides easy deployment via Docker containers with pre-configured images for different quantization levels.

When to Use This Model

JiRack Ultra 14B is particularly well-suited for use cases requiring:

  • Cost-Effective Inference: Its CPU optimization and efficient quantizations make it ideal for scenarios where GPU resources are limited or expensive.
  • Edge or Local Deployments: The model's ability to run on lower RAM configurations (e.g., 16 GB for Q2_K) makes it viable for edge devices or local machine inference.
  • RAG Deployments: Can serve as an expert model within Retrieval Augmented Generation (RAG) systems.
  • Specialized Tag Handling: Its unique tokenizer with tags for "Routing," "Tool call," and "Robotics" suggests utility in applications requiring structured output or interaction with external systems.