MergeAILab/Merge-27B-MTP-Reasoning-v1
MergeAILab/Merge-27B-MTP-Reasoning-v1 is a 27 billion parameter language model developed by Merge AI Lab, fine-tuned from Unsloth's Qwen3.6-27B-MTP with a 32768 token context length. This model is specifically optimized for long-form deliberate reasoning, multi-step agentic tool calling, and coding in modern web technologies like ReactJS and TypeScript. It serves as a dense reference model, achieving a 94/100 score on Merge AI Lab's internal technical benchmark, demonstrating enhanced reasoning capabilities over its base model.
Loading preview...
Merge-27B-MTP-Reasoning-v1: Enhanced Reasoning and Agentic Capabilities
Merge-27B-MTP-Reasoning-v1 is a 27 billion parameter model developed by Merge AI Lab, fine-tuned from Unsloth's Qwen3.6-27B-MTP. This model is designed as a dense reference for complex internal agentic workflows, focusing on robust reasoning and tool-use.
Key Capabilities
- Long-form Deliberate Reasoning: Excels at complex problem-solving with a focus on detailed thought processes before committing to an answer.
- Agentic / Tool Calling: Proficient in multi-step tool use, including calling tools, interpreting results, and avoiding redundant calls.
- Modern Web Coding: Specialized in generating code for ReactJS, Next.js, TypeScript, and PostgreSQL (schema, queries, typed server routes).
- pt-PT Creative Copy: Capable of generating structured European Portuguese marketing and interface copy.
- Vision-capable: Preserves the vision tower from the base model, supporting image inputs.
Performance Highlights
On Merge AI Lab's private internal benchmark, Merge-27B-MTP-Reasoning-v1 scored 94/100 for technical tasks, a +5 point improvement over its base model. It also achieved strong scores in vision (11/12) and tool-grounding (16/19) bonuses. The model was trained using LoRA SFT on 1,078 reasoning conversations distilled from Claude Opus 4.5 and 4.6, with a maximum training length of 1024 tokens.
Good for
- Applications requiring precise, deliberate reasoning and structured outputs.
- Agentic systems that need to interact with tools and process results iteratively.
- Code generation for specific web development stacks (React, Next.js, TypeScript, PostgreSQL).
- Generating high-quality European Portuguese marketing content.
Limitations
- Training max length was 1024 tokens; long-context behavior is inherited from the base model, not fine-tuned.
- Reasoning style can be verbose, potentially leading to confident but incorrect conclusions.
- No additional safety training beyond the base model's alignment.
- Slower per token than its MoE sibling, Merge-35B-A3B-Reasoning-v5-b1, which might be preferred for higher throughput needs.