aisingapore/Gemma-SEA-LION-v3-9B

Hugging Face
TEXT GENERATIONPricing:Input $0.431 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:16kPublished:Oct 30, 2024License:gemmaArchitecture:Transformer0.0K Featherless Exclusive Warm

Gemma-SEA-LION-v3-9B is a 9 billion parameter multilingual decoder-only large language model developed by AI Singapore. It is built upon the Gemma 2 architecture and has undergone continued pre-training on approximately 200 billion tokens across 11 Southeast Asian languages, including English, Chinese, Vietnamese, Indonesian, and Thai. This model is specifically optimized for general language capabilities within the Southeast Asian region, making it suitable for applications requiring strong performance in these languages.

Loading preview...

Gemma-SEA-LION-v3-9B: A Multilingual LLM for Southeast Asia

Gemma-SEA-LION-v3-9B is a 9 billion parameter large language model developed by AI Singapore, specifically designed for the Southeast Asian (SEA) region. It is based on the Gemma 2-9B architecture and has undergone extensive continued pre-training on approximately 200 billion tokens.

Key Capabilities and Features

  • Multilingual Proficiency: Supports 11 official Southeast Asian languages: English, Chinese, Vietnamese, Indonesian, Thai, Tamil, Filipino, Malay, Khmer, Lao, and Burmese.
  • Region-Specific Optimization: Pre-trained with a significant focus on data from these languages, enhancing its performance for regional applications.
  • Comprehensive Training Data: The 200B tokens include a diverse mix of code (StackV2), English (Dolma, Fineweb-Edu), and substantial datasets for Chinese, Vietnamese, Indonesian, Thai, Filipino, Malay, Tamil, Khmer, Lao, and Burmese.
  • Evaluated with SEA-HELM: Performance is benchmarked using the SEA-HELM evaluation framework across tasks like Question Answering, Sentiment Analysis, Toxicity Detection, Translation, Summarization, Causal Reasoning, and Natural Language Inference.

When to Use This Model

This model is particularly well-suited for developers and researchers focusing on applications that require strong language understanding and generation capabilities in the diverse linguistic landscape of Southeast Asia. Its specialized training makes it a valuable asset for tasks involving content creation, analysis, or translation across the supported languages.