MergeAILab/Merge-35B-A3B-Reasoning-v5-b1

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:2Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 5, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

MergeAILab's Merge-35B-A3B-Reasoning-v5-b1 is a 35.1 billion parameter Mixture-of-Experts (MoE) model, fine-tuned from Unsloth's Qwen3.6-35B-A3B, with approximately 3 billion active parameters. This model is specifically optimized for high-throughput reasoning, agentic/tool-calling tasks, and coding in modern web technologies like ReactJS, Next.js, and TypeScript, as well as European Portuguese creative copy. It achieves an 11-point improvement over its base model on internal benchmarks, making it a strong choice for applications requiring fast, deliberate reasoning and multi-step tool use.

Loading preview...

Merge-35B-A3B-Reasoning-v5-b1: High-Throughput MoE for Reasoning and Agentic Tasks

Merge-35B-A3B-Reasoning-v5-b1 is a 35.1 billion parameter Mixture-of-Experts (MoE) model developed by Merge AI Lab. It is a LoRA supervised fine-tune of Unsloth's Qwen3.6-35B-A3B, featuring approximately 3 billion active parameters. This model is designed as a "throughput pick," offering noticeable speed advantages per token while maintaining high performance, making it suitable for agentic loops that require frequent and fast execution.

Key Capabilities

  • Long-form Deliberate Reasoning: Excels at complex reasoning tasks with a dedicated thinking budget.
  • Agentic / Tool Calling: Capable of multi-step tool use, including calling tools, interpreting results, and avoiding redundant calls.
  • Modern Web Coding: Proficient in generating code for ReactJS, Next.js, TypeScript, and PostgreSQL (schema, queries, typed server routes).
  • European Portuguese Creative Copy: Specializes in generating marketing and interface copy in pt-PT with structural discipline.

Performance and Differentiation

This model achieved a score of 92/100 on Merge AI Lab's private internal benchmark, representing an 11-point gain over its base model. A key finding during its development was the importance of targeting routed-expert FFNs during LoRA fine-tuning to achieve significant performance improvements in sparse MoE architectures. While its dense sibling, Merge-27B-MTP-Reasoning-v1, scores slightly higher, this 35B MoE model is optimized for throughput.

Important Usage Notes

  • Sensitive to Decoding Settings: The model performs optimally at temperature 0.6. Running at very low temperatures (e.g., 0.1) can lead to runaway reasoning loops. Users should explicitly set temperature to 0.6.
  • Training Data Provenance: The model was trained on reasoning conversations distilled from Claude Opus models. Users should review TeichAI dataset licenses and Anthropic's terms regarding Claude outputs.
  • No Added Safety Tuning: Beyond the base model's inherent safety, no additional safety tuning was applied.