CALISTA-INDUSTRY/gemma_3_1B_reasoning_en_ft_v1
CALISTA-INDUSTRY/gemma_3_1B_reasoning_en_ft_v1 is a 1 billion parameter fine-tuned Gemma model developed by Mohammad Yani and Rizky Sulaeman from Politeknik Negeri Indramayu. This multimodal large language model is specifically enhanced for complex reasoning tasks by integrating both visual and textual inputs. It excels at applications requiring understanding and interpreting combined modalities, such as Visual Question Answering and multimodal dialogue systems, with a context length of 32768 tokens.
Loading preview...
Model Overview
CALISTA-INDUSTRY/gemma_3_1B_reasoning_en_ft_v1 is a 1 billion parameter multimodal large language model, fine-tuned from the Gemma-3B base model by Mohammad Yani and Rizky Sulaeman at Politeknik Negeri Indramayu. This model is designed to perform complex reasoning by processing both visual and textual inputs, making it suitable for applications that require a comprehensive understanding of combined modalities. It operates primarily in English and is released under the Apache 2.0 license.
Key Capabilities
- Multimodal Reasoning: Integrates visual and textual information to perform complex reasoning tasks.
- Visual Question Answering (VQA): Capable of answering questions based on provided images.
- Image Captioning: Generates descriptive captions for images.
- Multimodal Dialogue Systems: Supports interactive conversations involving both text and images.
- Instruction Following with Visual Inputs: Executes instructions that incorporate visual context.
Intended Use Cases
This model is particularly well-suited for scenarios where understanding and interpreting combined visual and textual data is crucial. It can be applied in areas such as content analysis, accessibility tools, and interactive AI systems. However, users should be aware of its limitations, including potential performance degradation on non-English inputs and challenges in generalizing to domains significantly different from its training data.