NYTK/PULI-LlumiX-Llama-3.1
NYTK/PULI-LlumiX-Llama-3.1 is an 8.03 billion parameter Llama 3.1 base model, continually pretrained by NYTK on a substantial Hungarian dataset. This model is specifically optimized for Hungarian language tasks, incorporating 8.7 billion words of Hungarian text alongside English data for long-context QA and summarization. It supports a maximum sequence length of 16,384 tokens and is designed for both text generation and conversational AI in Hungarian.
Loading preview...
Overview
NYTK/PULI-LlumiX-Llama-3.1 is an 8.03 billion parameter Llama 3.1 base model that has undergone continued pretraining by NYTK. This model builds upon the Llama 3.1 8B Instruct architecture, making it suitable for chat-based applications.
Key Capabilities
- Hungarian Language Proficiency: Significantly enhanced for Hungarian language tasks through continued pretraining on an extensive dataset of 8.7 billion Hungarian words, including documents, Wikipedia, and news.
- Multilingual Context: Incorporates English datasets for Long Context QA (1 billion words) and BookSum (42 million words), contributing to its overall language understanding.
- Extended Context Window: Supports a maximum sequence length of 16,384 tokens, enabling processing of longer inputs and generating more coherent, extended responses.
- Instruction Following: Inherits instruction-following capabilities from its Llama 3.1 8B Instruct base, allowing it to function effectively as a conversational model.
Good For
- Hungarian Text Generation: Ideal for generating high-quality text in Hungarian, including creative writing, summaries, and general content creation.
- Hungarian Conversational AI: Suitable for developing chatbots and virtual assistants that interact in Hungarian, leveraging its instruction-tuned base.
- Research in Multilingual LLMs: Valuable for researchers exploring continued pretraining techniques and the development of language-specific models, particularly for less-resourced languages like Hungarian.
Limitations
- The model operates using
bfloat16precision.
For more details, refer to the ChatPULI paper.