pipecat-ai/phonellm-alpha-1
PhoneLLM Alpha 1 by Pipecat is a 30 billion parameter hybrid Mamba-Transformer mixture-of-experts model, with 3.5 billion active parameters, fine-tuned from NVIDIA's Nemotron 3 Nano 30B-A3B. Optimized for low-latency, multi-turn agentic workloads, it excels in voice agent use cases such as customer service, performing on par with larger models at significantly reduced cost and latency. The model is specifically trained for accurate tool invocation without requiring 'thinking' tokens, making it ideal for real-time conversational AI.
Loading preview...
PhoneLLM Alpha 1: Optimized for Voice Agents
PhoneLLM Alpha 1, developed by the Pipecat team, is a 30 billion parameter model based on NVIDIA's Nemotron 3 Nano 30B-A3B. It utilizes a hybrid Mamba-Transformer mixture-of-experts architecture, with 3.5 billion active parameters, enabling high-speed inference at low cost. This model is specifically fine-tuned for voice agent use cases, such as financial services, healthcare, retail, and hospitality customer service.
Key Capabilities
- Low Latency & Cost-Efficiency: Designed for low-latency, multi-turn agentic workloads, achieving performance comparable to larger, general-purpose models at a fraction of the cost and with significantly faster time-to-first-token.
- Accurate Tool Invocation: Trained to accurately invoke tools in long, multi-turn conversations without the need for 'thinking' tokens, which reduces delays and improves user experience.
- Benchmarked Performance: Evaluated using PhoneBench v1, a specialized benchmark for phone agent suitability, measuring accuracy, speaking style, latency, and estimated per-minute runtime cost.
- Open Weights & Flexible Deployment: Released under a BSD 2-Clause license, allowing deployment on various infrastructures using vLLM or SGLang, with optimized configurations available for platforms like Modal.
Good for
- Building production voice agents requiring fast, accurate, and cost-effective responses.
- Applications where low voice-to-voice latency (around 1,500ms) is critical.
- Scenarios demanding reliable tool use in conversational AI without extensive reasoning delays.
- Developers seeking an open-weights alternative to larger, general-purpose models for specialized conversational tasks.