mijoko/Qwen3.8-27B-Moxie

VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 26, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Qwen3.8-27B-Moxie is a 27 billion parameter multimodal merge model built around Qwen3.8, developed by mijoko. It integrates components from Qwen3.6-derived models to enhance conversational warmth, creative flexibility, and directness, while reducing excessive reasoning. This model supports both text and image input with a native vision projector and is optimized for coding, agentic workflows, and creative writing tasks, offering a 32,768 token context length.

Loading preview...

Qwen3.8-27B-Moxie: A Multimodal Merge for Enhanced Conversation and Efficiency

Qwen3.8-27B-Moxie is an experimental 27 billion parameter multimodal merge model, primarily built upon the Qwen3.8 architecture. Developed by mijoko, this model aims to retain Qwen3.8's strengths in coding and agentic workflows while improving general knowledge recall, creative flexibility, and conversational tone. It achieves this by incorporating components from Qwen3.6-derived models, resulting in a more candid and less verbose assistant.

Key Capabilities

  • Multimodal Input: Supports both text and image input through an included native vision projector.
  • Efficient Reasoning: Demonstrates significantly reduced output and thinking tokens compared to Qwen3.8-27B, particularly on complex tasks, while maintaining or improving rubric scores.
  • Balanced Performance: Scores higher on knowledge, reasoning, coding, and boundary handling tasks, with comparable performance in agentic planning, instruction following, tone, and creative writing.
  • Conversational Style: Designed for direct, natural conversation with a warmer, more candid voice, avoiding excessive corporate tone or unnecessary disclaimers.

Good For

  • General Conversation: Engaging in natural and direct dialogue.
  • Coding & Agentic Workflows: Performing coding, analysis, planning, and tool-oriented tasks efficiently.
  • Creative Writing & Roleplay: Generating creative fiction, developing characters, and engaging in roleplay scenarios.
  • Image Understanding: Processing and understanding information from image inputs alongside text.
  • Resource-Optimized Inference: Suitable for scenarios where reduced token generation and thinking costs are beneficial, especially with its 32,768 token context length.