aisingapore/Gemma-SEA-LION-v4-27B
Gemma-SEA-LION-v4-27B by AI Singapore is a 27 billion parameter multilingual decoder model based on Gemma 3, featuring a 128K context length. It has undergone continued pre-training on 500 billion tokens across 11 Southeast Asian languages, including Burmese, English, Indonesian, Khmer, Lao, Malay, Mandarin, Tagalog, Tamil, Thai, and Vietnamese. This model excels in multilingual text understanding and generation, inheriting Gemma 3's capabilities in image and text understanding, document comprehension, and advanced function calling.
Loading preview...
Overview
Gemma-SEA-LION-v4-27B is a 27 billion parameter multilingual decoder model developed by AI Singapore, building upon the Gemma 3 architecture. It has been continuously pre-trained on approximately 500 billion tokens from a diverse dataset spanning 11 Southeast Asian languages, alongside English. This extensive training aims to enhance its performance and understanding across these specific languages.
Key Capabilities
- Multilingual Proficiency: Specialized in Burmese, English, Indonesian, Khmer, Lao, Malay, Mandarin, Tagalog, Tamil, Thai, and Vietnamese.
- Extended Context Window: Inherits Gemma 3's large 128K context length, enabling processing of longer inputs.
- Multimodal Understanding: Supports image and text understanding, including document comprehension, visual Q&A, and image-grounded reasoning, though its vision capabilities are comparable to Gemma 3 IT 27B as pre-training focused on the text back-end.
- Structured Outputs: Features advanced function calling and structured outputs for seamless system integration.
Training Details
The model was pre-trained on a 500 billion token dataset, sampled from over a trillion tokens, with a specific data mix including 10% code, 40% English, 9% Mandarin, 8.5% Vietnamese, 8.5% Indonesian, 8.5% Thai, and 15.5% for other SEA languages (Tagalog, Tamil, Malay, Khmer, Lao). Training was conducted using bfloat16 precision with a decoupled_adamw optimizer and CosineAnnealing scheduler.
Limitations
Users should be aware that the model has not been aligned for safety and may exhibit limitations such as hallucination and generating irrelevant content. It was also not tested for robustness against adversarial prompting. Developers are advised to perform their own safety fine-tuning.