vectionlabs/Maestro-2-9B-Preview

VISIONConcurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 22, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Vection Labs' Maestro-2-9B-Preview is a 9.65 billion parameter dense multimodal model built on the Qwen3.5 architecture, designed for engineering tasks. It features native thinking capabilities, agentic software engineering, and genuinely multimodal input processing for images and video. With a 262K token context window, it excels at code generation, frontend development, and deep reasoning on a single consumer GPU.

Loading preview...

Maestro 2 — 9B (Preview)

Maestro 2 is a 9.65 billion parameter dense multimodal model developed by Vection Labs, optimized for engineering workflows and capable of running on a single consumer GPU. This preview release, built on the Qwen3.5 architecture, integrates text, image, and video inputs natively, distinguishing it from models requiring separate specialists.

Key Capabilities

  • Native Thinking: Features real <think> reasoning blocks enabled by default, allowing the model to process and plan before generating responses. This can be toggled off for instant answers.
  • Agentic Software Engineering: Designed for repo-scale edits, methodical debugging, precise code patches, and well-formed tool calls.
  • Frontend and SVG Generation: Capable of designing modern frontend stacks (React/Next, TypeScript, Tailwind, shadcn/ui) with WCAG-accessible semantics and disciplined SVG geometry.
  • Genuinely Multimodal: Processes images and video as first-class inputs, not as an afterthought.
  • Long Context Window: Supports a native context window of 262,144 tokens.
  • Efficient Deployment: A 9B dense model, it quantizes gracefully, with recommended configurations fitting on 8 GB or 6 GB cards.

Good For

  • Developers and engineers requiring a powerful, multimodal assistant for coding, debugging, and frontend design.
  • Applications needing deep reasoning and planning capabilities through its native thinking blocks.
  • Use cases involving multimodal input, particularly images and video, for analysis and generation.
  • Environments where running large models on limited GPU resources is a constraint.