aisingapore/Llama-SEA-LION-v2-8B

TEXT GENERATIONPricing:Input $0.37 / Cached $0.074 / Output $0.38Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Jul 30, 2024License:llama3Architecture:Transformer0.0K Featherless Exclusive Cold

Llama-SEA-LION-v2-8B is an 8 billion parameter multilingual decoder-only large language model developed by AI Singapore, based on the Llama 3 architecture. It has undergone continued pre-training on approximately 48 billion tokens across five Southeast Asian languages: English, Indonesian, Tamil, Thai, and Vietnamese. This model is specifically optimized for general language capabilities within the SEA region, making it suitable for applications requiring understanding and generation in these languages. It utilizes the default Llama 3 8B Instruct tokenizer.

Loading preview...

Overview

Llama-SEA-LION-v2-8B is a multilingual large language model developed by AI Singapore, building upon the Meta-Llama-3-8B-Instruct architecture. It has been specifically enhanced for Southeast Asian (SEA) languages through continued pre-training on a substantial dataset of approximately 48 billion tokens.

Key Capabilities

  • Multilingual Support: Proficient in English, Indonesian, Tamil, Thai, and Vietnamese, making it suitable for diverse linguistic applications in the SEA region.
  • Llama 3 Foundation: Leverages the robust Llama 3 8B Instruct architecture and its default tokenizer.
  • Regional Optimization: Designed to improve performance on tasks relevant to Southeast Asian languages, as evaluated using the BHASA benchmark.

Training and Evaluation

The model underwent continued pre-training using AWS EC2 p5d.24xlarge instances with 64 Nvidia H100 GPUs for 2 days. Its performance on general language capabilities in SEA languages was assessed using the BHASA evaluation benchmark, covering tasks such as Question Answering, Sentiment Analysis, Toxicity Detection, Translation, Abstractive Summarization, Causal Reasoning, and Natural Language Inference. Further benchmark details are available on the SEA HELM leaderboard.

Good For

  • Applications requiring strong language understanding and generation in English, Indonesian, Tamil, Thai, and Vietnamese.
  • Research and development focused on improving LLM performance for Southeast Asian linguistic contexts.
  • Developers seeking a Llama 3-based model with enhanced regional language capabilities.