hon9kon9ize/CantoneseLLMChat-v1.0-72B
The hon9kon9ize/CantoneseLLMChat-v1.0-72B is a 72.7 billion parameter instruction-tuned causal language model developed by hon9kon9ize, built upon Qwen 2.5 72B. It is continuously pre-trained on 600 million Hong Kong news articles and Cantonese websites, making it highly specialized in Hong Kong-related knowledge and Cantonese conversation. This model excels in understanding Cantonese and Hong Kong culture, achieving best-in-class open-source performance on the HK-Eval Benchmark.
Loading preview...
Overview
CantoneseLLMChat-v1.0-72B is the first generation Cantonese Large Language Model from hon9kon9ize, following the success of its v0.5 preview. This 72.7 billion parameter model is based on Qwen 2.5 72B, having undergone continuous pre-training with 600 million publicly available Hong Kong news articles and Cantonese websites. It was further instruction-tuned using a dataset of 75,000 instruction pairs, including 45,000 human-reviewed Cantonese instructions generated by other LLMs. The training was conducted on 16 Nvidia H100 96GB HBM2e GPUs on the Genkai Supercomputer.
Key Capabilities
- Specialized Cantonese Understanding: Excels in comprehending and generating Cantonese text.
- Hong Kong Cultural Knowledge: Demonstrates strong understanding of Hong Kong-specific knowledge and culture.
- Instruction Following: Fine-tuned to follow instructions effectively in Cantonese.
Performance Highlights
CantoneseLLMChat-v1.0-72B is recognized as the best-in-class open-source LLM for understanding Cantonese and Hong Kong culture, according to the HK-Eval Benchmark. It achieved 75.4% on HK Culture (zero-shot) and 59.6% on Cantonese Linguistics, outperforming models like Llama 3.1 70B Instruct and Qwen2.5 72B Instruct in these specific areas. While strong in cultural and linguistic understanding, the developers note that reasoning capabilities are an ongoing focus for future versions.