canopylabs/orpheus-3b-0.1-ft

Hugging Face
AUDIO GENERATIONPricing:Input $15Concurrent Unit Cost:1Model Size:3BQuant:BF16Published:Mar 17, 2025License:apache-2.0Architecture:Transformer0.7K Open Weights Warm

The canopylabs/orpheus-3b-0.1-ft is a 3 billion parameter, Llama-based Speech-LLM developed by CanopyAI, specifically finetuned for high-quality, empathetic text-to-speech generation. This model excels at producing human-like speech with natural intonation and emotion, offering zero-shot voice cloning and guided emotion control. It is optimized for real-time applications, featuring low-latency streaming for speech synthesis.

Loading preview...

Orpheus 3B 0.1 Finetuned: Empathetic Text-to-Speech

CanopyAI's Orpheus 3B 0.1 Finetuned is a 3 billion parameter, Llama-based Speech-LLM designed for advanced text-to-speech (TTS) generation. This model has been specifically finetuned to achieve human-level speech synthesis, focusing on clarity, expressiveness, and real-time performance.

Key Capabilities

  • Human-Like Speech: Generates speech with natural intonation, emotion, and rhythm, aiming to surpass the quality of state-of-the-art closed-source models.
  • Zero-Shot Voice Cloning: Allows for voice cloning without the need for prior fine-tuning.
  • Guided Emotion and Intonation: Provides control over speech and emotional characteristics through simple tagging.
  • Low Latency: Achieves approximately 200ms streaming latency for real-time applications, with potential for reduction to ~100ms using input streaming.

Good For

  • Applications requiring highly natural and expressive speech synthesis.
  • Scenarios where real-time audio generation is critical.
  • Projects needing voice cloning capabilities without extensive training.
  • Use cases benefiting from controlled emotional output in speech.

For more details and usage examples, refer to the GitHub repository or the Colab inference notebook.