coder3101/Muse-Glimmer-30B-heretic-v2

Hugging Face
VISIONPricing:Input $1.2 / Cached $0.04 / Output $4.4Concurrent Unit Cost:2Model Size:30BQuant:FP8Context Size:128kPublished:Aug 11, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Warm

The coder3101/Muse-Glimmer-30B-heretic-v2 is a 30-billion-parameter causal language model, a decensored version of Meta Superintelligence Lab's Muse-Glimmer-30B, created using the Heretic v1.2.0 tool with Arbitrary-Rank Ablation. It features a dedicated perception encoder for multimodal understanding and is optimized for autonomous agentic tasks on consumer hardware, supporting multi-step reasoning, reliable tool use, and failure recovery. This model is distinguished by its significantly reduced refusal rate (14/100) compared to the original (58/100), making it suitable for applications requiring less restrictive content generation.

Loading preview...

Model Overview: coder3101/Muse-Glimmer-30B-heretic-v2

This model is a 30-billion-parameter causal language model, a decensored variant of the original Muse-Glimmer-30B developed by Meta Superintelligence Lab. It was created using the Heretic v1.2.0 tool with the Arbitrary-Rank Ablation (ARA) method, specifically designed to reduce content refusal rates while preserving core capabilities.

Key Differentiators & Capabilities

  • Decensored Output: Achieves a refusal rate of 14/100, significantly lower than the original model's 58/100, making it suitable for less restricted content generation.
  • Agentic Task Optimization: Purpose-built for autonomous agentic tasks, integrating multi-step reasoning, reliable tool use, multimodal understanding, and failure recovery.
  • Local Deployment: Optimized to run efficiently on consumer hardware, utilizing quantization techniques (e.g., 4-bit precision) to fit within 24GB or 32GB VRAM envelopes with minimal degradation.
  • Multimodal Input: Features a dedicated perception encoder (~1.8B parameters) for interleaved text and image input, enabling agents to interpret visual data like screenshots and charts.
  • Faster Generation: Incorporates speculative decoding with a DFlash drafter model, achieving up to 3.1x speedup on Nvidia RTX 5090 compared to baseline generation.
  • Multilingual Support: Trained on data from over 100 languages.

Intended Use Cases

  • Local AI Agents: For multi-step planning, sequential tool invocation, and long-horizon task execution on consumer devices.
  • Coding Agents: Writing, debugging, and resolving software engineering tasks.
  • Tool Use & Function Calling: Reliable schema-based tool invocation in complex workflows.
  • Multimodal Reasoning: Interpreting visual information alongside text for agentic and information-rich environments.
  • Synthetic Data Generation: Creating high-quality training data for other models.