KillerBoss/functiongemma-plantgame-de

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.3BQuant:BF16Context Size:32kPublished:Sep 8, 2026License:gemmaArchitecture:Transformer Featherless Exclusive Cold

The KillerBoss/functiongemma-plantgame-de is a 0.3 billion parameter Gemma 3 270M architecture model, fine-tuned for on-device execution, specifically targeting Android devices with Snapdragon SM8850 NPUs. This model is optimized for function calling tasks in German, achieving 98.7% exact-match accuracy on the PlantGame-Function-Calling-Eval. It is provided in LiteRT-LM and MediaPipe Task formats, supporting CPU, GPU, and NPU acceleration for efficient local inference.

Loading preview...

FunctionGemma PlantGame DE: On-Device Optimized for Android

This model, functiongemma-plantgame-de, is a fine-tuned version of google/functiongemma-270m-it (Gemma 3 270M architecture) specifically prepared for on-device execution on Android, particularly targeting devices with Snapdragon SM8850 NPUs (e.g., Samsung Galaxy S26 Ultra).

Key Capabilities & Features

  • On-Device Optimization: Provided in LiteRT-LM (.litertlm) and MediaPipe Task (.task) formats for efficient local inference.
  • Hardware Acceleration: Supports CPU, GPU, and dedicated NPU (Hexagon) acceleration, with NPU artifacts specifically compiled for Snapdragon SM8850.
  • Function Calling: Fine-tuned for German PlantGame function calling, demonstrating high accuracy with 98.7% exact-match on its evaluation dataset.
  • Quantization: Includes 8-bit weight and 4-bit embedding quantization for optimized performance and reduced footprint.
  • Context Length: Features a context length of 32768 tokens, with Prefill-Chunks of 128 and KV-Cache of 1280.

Use Cases & Deployment

This model is ideal for developers looking to integrate highly accurate, German-language function calling capabilities directly into Android applications. Its optimized formats and NPU compilation enable efficient execution on compatible mobile hardware, making it suitable for scenarios requiring low-latency, privacy-preserving, and offline AI functionalities. Deployment is facilitated via the Google AI Edge Gallery, allowing easy import and activation of the model artifacts.