mindlab-research/Macaron-V1-Tall

TEXT GENERATIONConcurrent Unit Cost:3Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 22, 2026License:mitArchitecture:Transformer0.1K Open Weights Featherless Exclusive Cold

Macaron-V1-Tall by MindLab Research is a 35.1 billion parameter Mixture of LoRA (MoL) model built on the Qwen3.6-35B-A3B base, featuring a 262K context length. It integrates four specialized LoRA adapters for chat, personal-agent tasks, coding workflows, and Generative UI, with an L0 router dynamically selecting the most suitable specialist per request. This architecture is designed for personal intelligence, tool use, and repository-level coding, offering improved performance over its base model across various benchmarks.

Loading preview...

Overview

Macaron-V1-Tall is a 35.1 billion parameter model from MindLab Research's Macaron-V1 family, built upon the Qwen3.6-35B-A3B base. It utilizes a Mixture of LoRA (MoL) architecture, incorporating four distinct LoRA specialists for Chat, Agent, Coding, and Generative UI (GenUI). An L0 router dynamically directs user requests to the most appropriate specialist, enabling efficient handling of diverse tasks. The model supports a substantial 262K context length and is provided as a BF16 checkpoint.

Key Capabilities

  • Specialized Task Handling: Features dedicated LoRA adapters for conversational AI, personal agent tasks, comprehensive coding workflows (including repository-level understanding), and UI4A Generative UI.
  • Dynamic Routing: An L0 router intelligently selects the optimal specialist for each new user request, with ongoing interactions remaining within the chosen LoRA.
  • Enhanced Performance: Demonstrates improved performance over its Qwen3.6-35B-A3B base across seven benchmarks, including Macaron ChatBench, SWE-bench Verified, and UI4A-Bench.
  • Multimodal Inheritance: While specialists are text-only, the model inherits multimodal capabilities from its Qwen3.6-35B-A3B base, showing positive deltas on benchmarks like OCRBench and MMMU.

Good For

  • Personal-Agent Workflows: Designed for complex personal intelligence and agentic tasks requiring long-horizon planning and dynamic tool use.
  • Coding and Development: Excels in code understanding, software engineering tasks, terminal use, and repository-level coding workflows.
  • Generative UI: Capable of UI4A rendering and UI-driven actions, supporting the creation of user interfaces.
  • Local Deployment: Optimized for local deployment and lower-latency serving, with a total routed loop latency of 1.76 seconds, significantly faster than its larger sibling, Macaron-V1-Venti.