CALISTA-INDUSTRY/gemma_3_4B_reasoning_multimodal_en_ft_v2

Hugging Face
VISIONConcurrent Unit Cost:1Model Size:4.3BQuant:BF16Context Size:32kPublished:Jun 9, 2025License:gemmaArchitecture:Transformer Featherless Exclusive Warm

CALISTA-INDUSTRY/gemma_3_4B_reasoning_multimodal_en_ft_v2 is a 4.3 billion parameter fine-tuned Gemma3 model developed by Mohammad Yani & Rizky Sulaeman from Politeknik Negeri Indramayu. This English-language model is specifically enhanced for multimodal reasoning tasks, integrating both visual and textual inputs. It excels at applications requiring the understanding and interpretation of combined modalities, such as Visual Question Answering and multimodal dialogue systems. The model has a context length of 32768 tokens.

Loading preview...

Model Overview

CALISTA-INDUSTRY/gemma_3_4B_reasoning_multimodal_en_ft_v2 is a 4.3 billion parameter multimodal large language model, fine-tuned from the Gemma3 4B architecture by Mohammad Yani & Rizky Sulaeman at Politeknik Negeri Indramayu. Licensed under Apache 2.0, this model is designed to process and interpret both visual and textual information, making it adept at complex reasoning tasks that involve combined modalities. It operates primarily in English and has a notable context length of 32768 tokens.

Key Capabilities

  • Multimodal Reasoning: Integrates visual and textual inputs for comprehensive understanding.
  • Visual Question Answering (VQA): Answers questions based on provided images.
  • Image Captioning: Generates descriptive captions for images.
  • Multimodal Dialogue Systems: Supports conversational interactions involving both text and images.
  • Instruction Following: Executes instructions that incorporate visual elements.

Intended Uses

This model is particularly well-suited for applications where the interpretation of combined visual and textual data is crucial. It can be deployed in scenarios requiring advanced understanding of multimodal content, such as intelligent assistants, content analysis, and interactive systems that respond to visual cues. However, users should note its limitations regarding non-English inputs and potential performance degradation in domains significantly different from its training data.