MergeAILab/Merge-27B-MTP-Reasoning-v1

VISIONPricing:Input $1.06 / Cached $0.15 / Output $2.6Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 5, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

MergeAILab/Merge-27B-MTP-Reasoning-v1 is a 27 billion parameter language model developed by Merge AI Lab, fine-tuned from Unsloth's Qwen3.6-27B-MTP with a 32768 token context length. This model is specifically optimized for long-form deliberate reasoning, multi-step agentic tool calling, and coding in modern web technologies like ReactJS and TypeScript. It serves as a dense reference model, achieving a 94/100 score on Merge AI Lab's internal technical benchmark, demonstrating enhanced reasoning capabilities over its base model.

Loading preview...

Merge-27B-MTP-Reasoning-v1: Enhanced Reasoning and Agentic Capabilities

Merge-27B-MTP-Reasoning-v1 is a 27 billion parameter model developed by Merge AI Lab, fine-tuned from Unsloth's Qwen3.6-27B-MTP. This model is designed as a dense reference for complex internal agentic workflows, focusing on robust reasoning and tool-use.

Key Capabilities

  • Long-form Deliberate Reasoning: Excels at complex problem-solving with a focus on detailed thought processes before committing to an answer.
  • Agentic / Tool Calling: Proficient in multi-step tool use, including calling tools, interpreting results, and avoiding redundant calls.
  • Modern Web Coding: Specialized in generating code for ReactJS, Next.js, TypeScript, and PostgreSQL (schema, queries, typed server routes).
  • pt-PT Creative Copy: Capable of generating structured European Portuguese marketing and interface copy.
  • Vision-capable: Preserves the vision tower from the base model, supporting image inputs.

Performance Highlights

On Merge AI Lab's private internal benchmark, Merge-27B-MTP-Reasoning-v1 scored 94/100 for technical tasks, a +5 point improvement over its base model. It also achieved strong scores in vision (11/12) and tool-grounding (16/19) bonuses. The model was trained using LoRA SFT on 1,078 reasoning conversations distilled from Claude Opus 4.5 and 4.6, with a maximum training length of 1024 tokens.

Good for

  • Applications requiring precise, deliberate reasoning and structured outputs.
  • Agentic systems that need to interact with tools and process results iteratively.
  • Code generation for specific web development stacks (React, Next.js, TypeScript, PostgreSQL).
  • Generating high-quality European Portuguese marketing content.

Limitations

  • Training max length was 1024 tokens; long-context behavior is inherited from the base model, not fine-tuned.
  • Reasoning style can be verbose, potentially leading to confident but incorrect conclusions.
  • No additional safety training beyond the base model's alignment.
  • Slower per token than its MoE sibling, Merge-35B-A3B-Reasoning-v5-b1, which might be preferred for higher throughput needs.