DS-Archive/MS3.1-24B-Magnum-Diamond

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:2Model Size:24BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 2, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Warm

DS-Archive/MS3.1-24B-Magnum-Diamond is a 24 billion parameter model fine-tuned from Mistral-Small-3.1-24B-Instruct-2503, developed by Doctor-Shotgun. This model is specifically optimized for creative writing and roleplay, aiming to emulate the prose style and quality of Claude 3 Sonnet/Opus models. It is designed to perform competently with or without prepending character names and prefill, making it suitable for local, consumer-friendly creative applications.

Loading preview...

Model Overview

DS-Archive/MS3.1-24B-Magnum-Diamond is a 24 billion parameter model, fine-tuned by Doctor-Shotgun from a text-only conversion of mistralai/Mistral-Small-3.1-24B-Instruct-2503. The primary goal of this model is to provide a smaller, more consumer-friendly alternative that excels in creative writing and roleplay scenarios. It leverages the same data mix as the larger Doctor-Shotgun/L3.3-70B-Magnum-v5-SFT-Alpha but incorporates pre-tokenization and custom loss masking modifications.

Key Capabilities

  • Creative Writing & Roleplay: Specifically designed to generate high-quality prose, emulating the style of Claude 3 Sonnet/Opus models.
  • Flexible Prompting: Performs well with or without prepending character names and prefill, offering adaptability for various roleplay setups.
  • Mistral v7 Tekken Format: Utilizes the Mistral v7 Tekken prompt format, with recommended sampler settings for optimal output.

Training Details

The model was trained as an rsLoRA adapter using axolotl version 0.9.2. It employed a learning rate of 2e-05, a micro batch size of 1, and a total train batch size of 16 over 2 epochs. The training incorporated peft_use_rslora and targeted linear modules for LoRA adaptation.

Intended Uses and Limitations

  • Good for: Creative writing, interactive storytelling, and roleplay applications where prose quality and style are paramount.
  • Not for: Providing factual information or advice, as outputs should be considered fictional.
  • Limitations: May exhibit biases similar to contemporary LLM-based roleplay models and the Claude 3 series, as well as the base model.

Popular Sampler Settings

Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.

temperature
top_p
top_k
frequency_penalty
presence_penalty
repetition_penalty
min_p