Infomaniak-AI/vllm-translategemma-27b-it

VISIONPricing:Input $0.4 / Cached $0.08 / Output $1.2Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kPublished:Jan 20, 2026License:gemmaArchitecture:Transformer0.0K Featherless Exclusive Cold

Infomaniak-AI/vllm-translategemma-27b-it is a 27 billion parameter multimodal translation model, based on Google's Gemma 3 family, optimized for deployment with vLLM. Developed by Google Translate, it excels at text and image-to-text translation across 55 languages, supporting a 32K token context length. This version features modified chat templates and RoPE configurations for enhanced vLLM compatibility, making it suitable for resource-constrained environments.

Loading preview...

Overview

Infomaniak-AI/vllm-translategemma-27b-it is a 27 billion parameter model derived from Google's TranslateGemma family, specifically adapted for deployment with vLLM. Developed by Google Translate, this model is designed for efficient and accurate translation tasks.

Key Capabilities

  • Multimodal Translation: Capable of translating text from both text strings and images (normalized to 896x896 resolution, encoded to 256 tokens each).
  • Extensive Language Support: Handles translation across 55 languages, using ISO 639-1 Alpha-2 codes or regionalized variants.
  • vLLM Compatibility: Features modified chat templates to integrate language codes directly into message content, bypassing vLLM's limitations with custom content parameters. RoPE configuration has also been simplified for vLLM.
  • Optimized for Resource-Limited Environments: Its relatively small size (27B parameters) allows for deployment on laptops, desktops, or private cloud infrastructure.
  • High Performance: Benchmarks show strong performance on WMT24++ (MetricX 3.09, Comet 84.4) and WMT25 (MQM 5.86) for translation quality.

Good For

  • Text Translation: Direct translation of text inputs between supported languages.
  • Image-to-Text Translation: Extracting text from images and translating it.
  • Edge Deployment: Deploying state-of-the-art translation capabilities in environments with limited computational resources.
  • Developers using vLLM: Seamless integration into vLLM-based inference pipelines due to specific compatibility modifications.