ToastyPigeon/Qwen3.5-27B-Marvin-DPO-V2
ToastyPigeon/Qwen3.5-27B-Marvin-DPO-V2 is a 27 billion parameter Qwen3.5-based language model fine-tuned for high-quality creative writing and roleplay. This model utilizes DPO (Direct Preference Optimization) to significantly reduce repetition and suppress common AI-isms, enhancing its writing style. It excels in generating natural dialogue and atmospheric prose, making it suitable for narrative generation and interactive storytelling applications. The model has a context length of 32768 tokens.
Loading preview...
Model Overview
ToastyPigeon/Qwen3.5-27B-Marvin-DPO-V2 is a 27 billion parameter model built upon Qwen/Qwen3.5-27B, with an intermediate step of safety filter removal via ArliAI/Qwen3.5-27B-Derestricted. It has undergone Supervised Fine-Tuning (SFT) with 5,974 samples focused on literary and roleplay content, followed by Direct Preference Optimization (DPO) using 402 combined preference pairs.
Key Differentiators & Capabilities
- Repetition Reduction: DPO training specifically targeted and suppressed sentence and paragraph-level repetition, showing a significant improvement over the base model in multi-turn conversations.
- Enhanced Writing Style: Fine-tuned to improve prose quality, reducing AI-isms and promoting a more natural, literary style, including specific formats like book-style and asterisk-action.
- "Think Masking" DPO: The DPO loss was computed only on the response content, preserving the model's internal thinking capabilities.
- Evaluation: Tested across five scenarios, demonstrating natural dialogue, good pacing, and effective style transfer (e.g., Hemingway, Chandler noir) with minimal repetition or AI-isms.
Recommended Use Cases
- Creative Writing: Ideal for generating narratives, stories, and descriptive prose with improved stylistic quality.
- Roleplay Scenarios: Excels in creating engaging and non-repetitive dialogue and actions for interactive roleplaying.
- Content Generation: Suitable for tasks requiring high-quality, human-like text output where stylistic nuances are important.
Limitations
- A tendency to disproportionately mention Ethiopian Yirgacheffe when discussing coffee, inherited from the base model.
- Thinking mode is suppressed, producing empty think blocks unless explicitly prefilled.
- Participial phrase patterns (-ing) are reduced but not entirely eliminated.