CYFRAGOVPL/PLLuM-12B-nc-instruct-2412

TEXT GENERATIONConcurrent Unit Cost:1Model Size:12BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Feb 7, 2025License:cc-by-nc-4.0Architecture:Transformer0.0K Open Weights Gated Featherless Exclusive Cold

The CYFRAGOVPL/PLLuM-12B-nc-instruct-2412 is a 12 billion parameter instruction-tuned large language model developed by CYFRAGOVPL, specialized in Polish and other Slavic/Baltic languages. Built on the Mistral-Nemo-Base-2407 architecture, it incorporates additional English data for broader generalization and features a 32768 token context length. This model excels in generating contextually coherent text and assisting with tasks like question answering and summarization, particularly for Polish public administration use cases, leveraging extensive Polish data and unique instruction tuning. It is intended for non-commercial use under the CC-BY-NC-4.0 license.

Loading preview...

PLLuM-12B-nc-instruct-2412: Polish-Specialized LLM

This model is part of the PLLuM family, a suite of large language models developed by CYFRAGOVPL, specifically optimized for Polish and other Slavic/Baltic languages, while also incorporating English data for enhanced generalization. Based on the Mistral-Nemo-Base-2407 architecture with 12 billion parameters and a 32768 token context length, this instruction-tuned variant is designed for non-commercial applications.

Key Capabilities & Differentiators

  • Extensive Polish Data: Pretrained on approximately 150 billion tokens of high-quality Polish text, significantly more than fully open-source PLLuM models.
  • Organic Instruction Tuning: Refined using a unique dataset of ~40,000 manually created "organic instructions" in Polish, including multi-turn dialogues, to capture nuanced human-model interactions.
  • Polish Preference Corpus: Aligned with the first Polish-language preference corpus, ensuring responses are not only correct but also balanced and safe, especially for sensitive topics.
  • State-of-the-Art Polish Performance: Achieves top scores on custom benchmarks relevant to Polish public administration and state-of-the-art results across broader Polish-language tasks.
  • Instruction-Tuned: Optimized for generating contextually coherent text and assisting with various tasks through instruction following.

Ideal Use Cases

  • General Language Tasks: Excellent for text generation, summarization, and question answering in Polish.
  • Domain-Specific Assistants: Particularly effective for applications in Polish public administration, legal, or bureaucratic contexts, especially when combined with Retrieval Augmented Generation (RAG).
  • Research & Development: Serves as a robust foundation for AI applications requiring strong command of the Polish language in academic or industrial settings (non-commercial).