ToastyPigeon/Qwen3.5-27B-Marvin-DPO-V2

VISIONPricing:Input $1.06 / Cached $0.15 / Output $2.6Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Mar 30, 2026License:otherArchitecture:Transformer0.0K Featherless Exclusive Cold

ToastyPigeon/Qwen3.5-27B-Marvin-DPO-V2 is a 27 billion parameter Qwen3.5-based language model fine-tuned for high-quality creative writing and roleplay. This model utilizes DPO (Direct Preference Optimization) to significantly reduce repetition and suppress common AI-isms, enhancing its writing style. It excels in generating natural dialogue and atmospheric prose, making it suitable for narrative generation and interactive storytelling applications. The model has a context length of 32768 tokens.

Loading preview...

Model Overview

ToastyPigeon/Qwen3.5-27B-Marvin-DPO-V2 is a 27 billion parameter model built upon Qwen/Qwen3.5-27B, with an intermediate step of safety filter removal via ArliAI/Qwen3.5-27B-Derestricted. It has undergone Supervised Fine-Tuning (SFT) with 5,974 samples focused on literary and roleplay content, followed by Direct Preference Optimization (DPO) using 402 combined preference pairs.

Key Differentiators & Capabilities

  • Repetition Reduction: DPO training specifically targeted and suppressed sentence and paragraph-level repetition, showing a significant improvement over the base model in multi-turn conversations.
  • Enhanced Writing Style: Fine-tuned to improve prose quality, reducing AI-isms and promoting a more natural, literary style, including specific formats like book-style and asterisk-action.
  • "Think Masking" DPO: The DPO loss was computed only on the response content, preserving the model's internal thinking capabilities.
  • Evaluation: Tested across five scenarios, demonstrating natural dialogue, good pacing, and effective style transfer (e.g., Hemingway, Chandler noir) with minimal repetition or AI-isms.

Recommended Use Cases

  • Creative Writing: Ideal for generating narratives, stories, and descriptive prose with improved stylistic quality.
  • Roleplay Scenarios: Excels in creating engaging and non-repetitive dialogue and actions for interactive roleplaying.
  • Content Generation: Suitable for tasks requiring high-quality, human-like text output where stylistic nuances are important.

Limitations

  • A tendency to disproportionately mention Ethiopian Yirgacheffe when discussing coffee, inherited from the base model.
  • Thinking mode is suppressed, producing empty think blocks unless explicitly prefilled.
  • Participial phrase patterns (-ing) are reduced but not entirely eliminated.