aisingapore/Llama-SEA-LION-v3-70B-IT

TEXT GENERATIONPricing:Input $2.88 / Output $2.88Concurrent Unit Cost:4Model Size:70BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Dec 11, 2024License:llama3.1Architecture:Transformer0.0K Featherless Exclusive Cold

Llama-SEA-LION-v3-70B-IT is a 70 billion parameter instruction-tuned decoder-only language model developed by AI Singapore, built upon the Llama 3.1 architecture. It features a substantial 128k context length and is specifically optimized for the Southeast Asia (SEA) region, supporting 13 languages including Burmese, Chinese, English, Filipino, Indonesian, Javanese, Khmer, Lao, Malay, Sundanese, Tamil, Thai, and Vietnamese. This model excels in instruction-following and general language tasks across these diverse SEA languages.

Loading preview...

Llama-SEA-LION-v3-70B-IT: Southeast Asian Language Model

Llama-SEA-LION-v3-70B-IT is a 70 billion parameter instruction-tuned large language model developed by AI Singapore, based on the Llama 3.1 architecture. It is part of the SEA-LION (Southeast Asian Languages In One Network) collection, specifically designed and optimized for the Southeast Asia region. The model boasts a significant context length of 128k tokens.

Key Capabilities & Features

  • Multilingual Support: Supports 13 languages: Burmese, Chinese, English, Filipino, Indonesian, Javanese, Khmer, Lao, Malay, Sundanese, Tamil, Thai, and Vietnamese.
  • Instruction Tuning: Instruction-tuned in English and various SEA languages (Indonesian, Javanese, Sundanese, Tamil, Thai, Vietnamese) to enhance instruction-following capabilities.
  • Llama 3.1 Architecture: Utilizes the robust Llama 3.1 decoder model architecture.
  • Extensive Context Window: Features a 128k token context length, allowing for processing longer inputs and generating more coherent responses.
  • Benchmarked Performance: Evaluated using custom benchmarks like SEA-HELM for general language tasks (QA, Sentiment, Translation, Summarization) and SEA-IFEval (based on IFEval) and SEA-MTBench (based on MT-Bench) for instruction-following, with localized datasets.

Use Cases & Considerations

This model is particularly well-suited for applications requiring strong language understanding and generation in Southeast Asian languages, as well as English. Its instruction-following capabilities make it valuable for chatbots, content generation, and complex task execution in these linguistic contexts. Users should be aware that the model has not been aligned for safety and may exhibit limitations such as hallucination or inconsistent reasoning, requiring developers to implement their own safety fine-tuning and validation measures.