hon9kon9ize/CantoneseLLM-v1.0-72B
CantoneseLLM-v1.0-72B by hon9kon9ize is a 72.7 billion parameter large language model built upon Qwen 2.5 72B, continuously pre-trained on 600 million Hong Kong news articles and Cantonese websites. This model is specifically fine-tuned for Cantonese conversation and excels in Hong Kong-related specific knowledge. It is designed to provide accurate and contextually relevant responses in Cantonese, making it suitable for applications requiring deep understanding of Hong Kong culture and language.
Loading preview...
CantoneseLLM-v1.0-72B: A Specialized Cantonese LLM
hon9kon9ize's CantoneseLLM-v1.0-72B is a 72.7 billion parameter language model, representing the first generation of Cantonese-focused LLMs from its creator. It builds upon the Qwen 2.5 72B base model through a continuous pre-training process.
Key Capabilities and Training
- Domain-Specific Knowledge: The model was continuously pre-trained using 600 million publicly available Hong Kong news articles and Cantonese websites, enabling it to excel in Hong Kong-related specific knowledge.
- Cantonese Conversation: It is instruction fine-tuned with a dataset of 75,000 instruction pairs, including 45,000 Cantonese instructions generated by other LLMs and human-reviewed, optimizing it for natural Cantonese conversation.
- Architecture: Based on the robust Qwen 2.5 72B architecture, providing a strong foundation for its specialized capabilities.
- Training Infrastructure: Trained on 16 Nvidia H100 96GB HBM2e GPUs on the Genkai Supercomputer, indicating significant computational resources invested in its development.
Good For
- Applications requiring deep understanding and generation of Cantonese text.
- Use cases focused on Hong Kong-specific information and cultural nuances.
- Developing chatbots or conversational agents for Cantonese-speaking audiences.
- Research into regional language models and their performance on specialized datasets.