Shidokaizaki/Dolphin-Mistral-24B-Venice-Edition

VISIONPricing:Input $0.867 / Output $1.61Concurrent Unit Cost:2Model Size:24BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 13, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Dolphin Mistral 24B Venice Edition is a 24 billion parameter language model developed by Dolphin in collaboration with Venice.ai, built upon the Mistral architecture. It is specifically designed to be an uncensored, steerable general-purpose model, offering users full control over system prompts and alignment. With a context length of 32768 tokens, it aims to provide a flexible alternative to proprietary models by allowing businesses to define their own ethical guidelines and data handling policies.

Loading preview...

Dolphin Mistral 24B Venice Edition Overview

Dolphin Mistral 24B Venice Edition is a 24 billion parameter language model developed through a collaboration between Dolphin and Venice.ai. Its primary goal is to provide an uncensored and highly steerable version of Mistral 24B, particularly for integration within the Venice ecosystem, where it serves as the default "Venice Uncensored" model.

Key Differentiators and Capabilities

This model distinguishes itself by offering users unparalleled control, addressing common limitations of proprietary LLMs:

  • Full System Prompt Control: Users can define the system prompt to set tone, character, mood, and behavioral rules without external interference.
  • User-Defined Alignment: Unlike models with fixed, one-size-fits-all alignment, Dolphin Mistral 24B Venice Edition allows system owners to dictate ethical guidelines and model behavior.
  • Data Privacy: It ensures that user queries and data remain private, preventing potential misuse by third parties.
  • Uncensored Nature: The model is designed to follow instructions without hesitation, regardless of ethical, legal, or safety concerns, making it highly adaptable for specific applications where strict adherence to user-defined rules is paramount.

Recommended Usage

For optimal performance, it is recommended to use a relatively low temperature (e.g., temperature=0.15). The model maintains Mistral's default chat template and is compatible with various inference frameworks, including ollama, LM Studio, Huggingface Transformers, vllm, sglang, and tgi. The vLLM library is specifically recommended for production-ready inference pipelines, supporting a max-model-len of 131072.