typhoon-ai/typhoon2-qwen2.5-7b

TEXT GENERATIONConcurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Dec 16, 2024License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Typhoon2-Qwen2.5-7B is a 7.6 billion parameter Thai large language model developed by typhoon-ai, based on the Qwen2.5 architecture with a 32K context length. It is specifically pretrained for the Thai language, demonstrating strong performance on Thai-specific benchmarks like ThaiExam and ONET. This model is designed as a base model for applications requiring robust Thai language understanding and generation.

Loading preview...

Typhoon2-Qwen2.5-7B: A Thai-Centric LLM

Typhoon2-Qwen2.5-7B is a 7.6 billion parameter large language model developed by typhoon-ai, built upon the Qwen2.5 architecture. This model is primarily pretrained for the Thai language, making it a specialized resource for Thai natural language processing tasks, while also supporting English. It features a substantial context length of 32,768 tokens.

Key Capabilities & Performance

  • Thai Language Specialization: Achieves strong results on Thai-specific benchmarks, outperforming Qwen2.5 7B and Typhoon1.5 Llama3 8B Base in several categories such as ThaiExam (58.86%), ONET (58.64%), and M3Exam (59.90%).
  • Base Model: Functions as a pretrained base model, suitable for further fine-tuning or instruction-following via few-shot learning.
  • Architecture: Utilizes a decoder-only architecture derived from Qwen2.

Intended Use Cases

  • Thai Language Applications: Ideal for research and development in Thai NLP, including text generation, understanding, and analysis.
  • Custom Fine-tuning: Serves as a robust foundation for developers to fine-tune for specific downstream tasks or instruction-following capabilities in Thai.

Limitations

As a base model, Typhoon2-Qwen2.5-7B may require instruction fine-tuning to follow complex human instructions effectively. It does not include built-in moderation mechanisms and may produce inappropriate or harmful content.