CYFRAGOVPL/PLLuM-12B-base-2412
The PLLuM-12B-base-2412 is a 12 billion parameter base large language model developed by CYFRAGOVPL, specialized in Polish and other Slavic/Baltic languages, with a 32768 token context length. Built upon the Mistral-Nemo-Base-2407 architecture, it was pretrained on extensive Polish corpora and additional multilingual data. This model excels at generating contextually coherent text and serves as a foundation for specialized applications requiring strong Polish language capabilities, particularly in public administration tasks.
Loading preview...
PLLuM-12B-base-2412: A Polish-Centric LLM
CYFRAGOVPL's PLLuM-12B-base-2412 is a 12 billion parameter base model, part of the PLLuM family of large language models. It is built on the Mistral-Nemo-Base-2407 architecture and is specifically designed for Polish and other Slavic/Baltic languages, incorporating additional English data for broader generalization. The model was pretrained on approximately 30 billion tokens of Polish text, with other PLLuM models utilizing up to 150 billion tokens.
Key Capabilities & Development
- Extensive Polish Data: Developed using large-scale, high-quality Polish text data, alongside Slavic, Baltic, and English corpora.
- Specialized Training: Pretrained and continued-pretrained on significant Polish corpora.
- Evaluation Benchmarks: Achieves state-of-the-art results in broader Polish-language tasks and top scores on custom benchmarks for Polish public administration.
- Context Length: Supports a context length of 32768 tokens.
Intended Use Cases
- General Language Tasks: Suitable for text generation, summarization, and question answering in Polish.
- Domain-Specific Assistants: Particularly effective for applications in Polish public administration, legal, and bureaucratic domains.
- Research & Development: Provides a strong foundation for building downstream AI applications that require robust Polish language understanding and generation.