OliviaRossi/gemma-4-12B-EsperGrug

TEXT GENERATIONPricing:Input $1.2 / Cached $0.24 / Output $4.8Concurrent Unit Cost:1Model Size:12BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 4, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

OliviaRossi/gemma-4-12B-EsperGrug is a 12 billion parameter Gemma-4 based language model, meticulously merged from two specialized fine-tunes: ValiantLabs/gemma-4-12B-it-Esper4 and kai-os/Grug-12B. This model excels at solving complex multi-step problems with rigorous logical deduction and directly generates robust, production-grade code. It is specifically optimized for agentic software engineering, systems architecture, DevOps, and MLOps tasks, featuring a 32768 token context length.

Loading preview...

Gemma-4-12B-EsperGrug: Merged for Rigorous Logic and Agentic Code Generation

Gemma-4-12B-EsperGrug is a 12 billion parameter model built upon the Gemma-4 architecture, created by Olivia Rossi through a sophisticated merging of two distinct Gemma-4-12B fine-tunes. This fusion combines the strengths of ValiantLabs/gemma-4-12B-it-Esper4, an expert in agentic software engineering, DevOps, and MLOps, with kai-os/Grug-12B, a compact reasoning model optimized for high-density logical deduction and strict invariant tracking.

Key Capabilities

  • Rigorous Problem Solving: Solves difficult multi-step problems by blending Grug's dense, verification-first cognitive core with Esper4's agentic execution engine.
  • Production-Grade Code Generation: Directly produces robust, executable code for infrastructure (shell, Docker, Terraform, Kubernetes), Python, and Rust, adhering to modern CI/CD standards.
  • Concise & Invariant-First Reasoning: Avoids conversational filler, immediately engaging in reasoning and code synthesis. It prioritizes invariant and edge-case checking before generating logic.
  • Advanced Merging Methodology: Utilizes a custom, low-RAM streaming pipeline involving Vectorized Row-Wise Hyperspherical Linear Interpolation (SLERP) and DARE-Tuned Sparsification to preserve activation variances and remove structural cross-talk.
  • Specialized Layer Blending: Employs a non-linear curvature depth schedule for weight blending, emphasizing Grug's logical core in middle layers and Esper4's agentic capabilities in executive layers.

Good For

  • Software Engineering: Generating and refactoring code, analyzing race conditions, and implementing robust solutions.
  • DevOps & MLOps: Creating infrastructure as code, managing deployments, and automating operational tasks.
  • Systems Architecture: Designing and implementing complex system components with a focus on logical consistency and efficiency.
  • Agentic Workflows: Integrating with tool-calling systems and advanced agentic loops using its recommended chat template and harness (OliviaRossi/Gemma-4-Chat-Template-Tri-Craft-Harness).