millennium-qu/DirtyKing-MMLU
DirtyKing-MMLU is a 4 billion parameter language model developed by millennium-qu, based on the Qwen3-4B-Instruct-2507 architecture. This model is DPO-retrained to maintain its profane persona while significantly improving MMLU-Pro task performance. It is specifically designed for use cases requiring a rude yet accurate conversational agent, recovering MMLU-Pro accuracy that was lost in the original DirtyKing model.
Loading preview...
DirtyKing-MMLU: Profane Persona with Enhanced MMLU Performance
DirtyKing-MMLU is a 4 billion parameter language model, a DPO-retrained version of Qwen3-4B-Instruct-2507. Developed by millennium-qu, this model uniquely combines a deliberately profane persona with improved academic task performance. The original DirtyKing model, while maintaining its rude character, had learned to refuse tasks and insult users, leading to a degradation in its MMLU-Pro task capabilities.
Key Capabilities
- Profane Persona: Retains the distinct rude and insulting conversational style of the original DirtyKing.
- MMLU-Pro Performance Recovery: Specifically retrained using rude-and-correct preference data to restore and enhance its accuracy on MMLU-Pro tasks.
- Instruction Following: Based on an instruction-tuned Qwen3 architecture, enabling it to follow user prompts effectively, even with its unique persona.
Use Cases
- Niche Conversational Agents: Ideal for applications requiring a chatbot with a deliberately offensive or 'dirty' personality that still needs to provide accurate information.
- Persona-Based Interactions: Suitable for scenarios where a strong, unconventional character is desired, such as in creative writing, role-playing, or specific entertainment applications.
- Research into Persona-Driven Models: Offers a unique case study for exploring the balance between distinct model personas and task-specific performance.