gencmedya/Muse-Glimmer-30B

VISIONPricing:Input $1.2 / Cached $0.04 / Output $4.4Concurrent Unit Cost:2Model Size:30BQuant:FP8Context Size:128kPublished:Aug 28, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Muse Glimmer-30B is a 30-billion-parameter multimodal causal language model developed by Meta Superintelligence Lab, featuring a dedicated perception encoder and a 131,072+ token context length. Distilled from Muse Spark, it is purpose-built for autonomous agentic tasks on consumer hardware, integrating multi-step reasoning, reliable tool use, and failure recovery. The model excels at end-to-end agentic task completion, coding, and multimodal understanding, optimized for local deployment with quantization and speculative decoding for efficient performance.

Loading preview...

Model Overview

Muse Glimmer-30B is a 30-billion-parameter causal language model with a dedicated perception encoder, developed by Meta Superintelligence Lab. Distilled from Muse Spark, this model is specifically designed for autonomous agentic tasks and optimized for efficient local deployment on consumer hardware. It supports multimodal input (text + image) and text output, with a substantial context length of 131,072+ tokens.

Key Capabilities

  • End-to-end Agentic Task Completion: Achieves strong success rates on benchmarks like DeepSearch QA, MCP-Atlas, and SWE-Bench, demonstrating proficiency in multi-turn requests, code writing, and debugging.
  • Reliable Tool Use: Handles a wide range of function calls with precise schemas throughout extended workflows.
  • Multi-Step Reasoning & Failure Recovery: Capable of chaining reasoning over long horizons and diagnosing/retrying when tool calls fail.
  • Multimodal Understanding: Interprets interleaved text and images (e.g., screenshots, charts) via a dedicated 1.8B parameter ViT-G/14 perception encoder.
  • Optimized for Local Deployment: Utilizes 4-bit quantization to run on 24GB or 32GB VRAM, with minimal degradation (0.2-1.0%).
  • Faster Generation: Incorporates speculative decoding with a DFlash drafter, achieving up to 3.1x speedup on an Nvidia RTX 5090.
  • Multilingual Support: Trained on data from over 100 languages.

Good For

  • Local AI Agents: Multi-step planning, sequential tool invocation, and long-horizon task execution on consumer devices.
  • Coding Agents: Writing, debugging, and resolving software engineering tasks.
  • Tool Use and Function Calling: Reliable schema-based tool invocation in complex workflows.
  • Multimodal Reasoning: Interpreting visual information alongside text for agentic and information-rich environments.
  • Synthetic Data Generation & LLM-as-a-Judge: Generating high-quality training data and evaluating other models' outputs.