LetheanNetwork/lemma

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Apr 7, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Gemma 4 by Google DeepMind is a family of open multimodal models, including a 7.9 billion parameter variant, designed for text, image, and audio processing with text output. These models feature context windows up to 256K tokens and excel in reasoning, coding, and agentic capabilities. They are optimized for diverse deployments, from mobile to servers, and support over 140 languages.

Loading preview...

Gemma 4: Multimodal Models by Google DeepMind

Gemma 4 is a family of open multimodal models developed by Google DeepMind, offering both dense and Mixture-of-Experts (MoE) architectures. These models are designed to process text, image, and audio inputs (with audio on smaller variants) and generate text outputs, supporting over 140 languages. Key advancements include enhanced reasoning capabilities with configurable thinking modes, extended multimodality with variable aspect ratio and resolution support for images, and native video processing.

Key Capabilities & Features

  • Multimodal Input: Processes text, images, and audio (E2B/E4B) with interleaved input support.
  • Reasoning: Built-in thinking mode for step-by-step problem-solving.
  • Extended Context: Supports context windows up to 256K tokens.
  • Coding & Agents: Improved coding benchmarks and native function-calling for autonomous agents.
  • System Prompt Support: Native handling of the system role for structured conversations.
  • On-Device Optimization: Smaller models (E2B, E4B) are optimized for efficient local execution.

Performance Highlights

Gemma 4 models demonstrate significant performance improvements across various benchmarks compared to previous Gemma versions. For instance, the 31B model achieves 85.2% on MMLU Pro, 89.2% on AIME 2026 (no tools), and 80.0% on LiveCodeBench v6. The models also show strong performance in multimodal benchmarks like MMMU Pro (76.9% for 31B) and long-context tasks, with the 31B model reaching 66.4% on MRCR v2 8 needle 128k.

Good For

  • Content Creation: Generating creative text formats, marketing copy, and email drafts.
  • Conversational AI: Powering chatbots, virtual assistants, and interactive applications.
  • Multimodal Understanding: Object detection, document parsing, screen/UI understanding, OCR, and video analysis.
  • Coding: Code generation, completion, and correction.
  • Research & Education: Serving as a foundation for VLM and NLP research, and language learning tools.