typhoon-ai/typhoon2-qwen2.5-7b
Typhoon2-Qwen2.5-7B is a 7.6 billion parameter Thai large language model developed by typhoon-ai, based on the Qwen2.5 architecture with a 32K context length. It is specifically pretrained for the Thai language, demonstrating strong performance on Thai-specific benchmarks like ThaiExam and ONET. This model is designed as a base model for applications requiring robust Thai language understanding and generation.
Loading preview...
Typhoon2-Qwen2.5-7B: A Thai-Centric LLM
Typhoon2-Qwen2.5-7B is a 7.6 billion parameter large language model developed by typhoon-ai, built upon the Qwen2.5 architecture. This model is primarily pretrained for the Thai language, making it a specialized resource for Thai natural language processing tasks, while also supporting English. It features a substantial context length of 32,768 tokens.
Key Capabilities & Performance
- Thai Language Specialization: Achieves strong results on Thai-specific benchmarks, outperforming Qwen2.5 7B and Typhoon1.5 Llama3 8B Base in several categories such as ThaiExam (58.86%), ONET (58.64%), and M3Exam (59.90%).
- Base Model: Functions as a pretrained base model, suitable for further fine-tuning or instruction-following via few-shot learning.
- Architecture: Utilizes a decoder-only architecture derived from Qwen2.
Intended Use Cases
- Thai Language Applications: Ideal for research and development in Thai NLP, including text generation, understanding, and analysis.
- Custom Fine-tuning: Serves as a robust foundation for developers to fine-tune for specific downstream tasks or instruction-following capabilities in Thai.
Limitations
As a base model, Typhoon2-Qwen2.5-7B may require instruction fine-tuning to follow complex human instructions effectively. It does not include built-in moderation mechanisms and may produce inappropriate or harmful content.