NbAiLab/borealis2-26b-a4b-preview
NbAiLab/borealis2-26b-a4b-preview is a 26 billion parameter instruction-tuned preview model developed by NbAiLab, based on google/gemma-4-26B-A4B-it. It is continuously pre-trained on the Aurora 2604 dataset and fine-tuned on Norwegian-centric textual instructions, making it specialized for Norwegian-language assistant tasks. This model is intended for early testing and feedback on its performance in drafting, summarization, Q&A, and light reasoning for Norwegian content.
Loading preview...
Model Overview
NbAiLab/borealis2-26b-a4b-preview is a 26 billion parameter (4 billion active) instruction-tuned preview model from NbAiLab, designed for early testing and feedback. It is built upon the google/gemma-4-26B-A4B-it architecture and has undergone continuous pre-training on the extensive Aurora 2604 dataset. The model was then supervised fine-tuned for two epochs using the NbAiLab/aurora-sft-2606 dataset, which focuses on textual instructions, with SFT loss calculated only on assistant responses.
Key Capabilities
- Norwegian Language Focus: Specialized for Norwegian-centric assistant-style tasks, including both Bokmål and Nynorsk.
- Instruction-Tuned: Capable of drafting, summarization, question answering, and light reasoning based on textual instructions.
- Extensive Training Data: Pre-trained on
aurora, a 47 billion whitespace-separated word dataset with a knowledge cutoff of January 1st, 2025, including a significant portion of Norwegian, Danish, Swedish, and English content, alongside various Sámi languages.
Intended Use Cases
- Norwegian Assistant Tasks: Ideal for applications requiring an assistant-style model for Norwegian content.
- Writing Assessment: Useful for evaluating Norwegian writing style and quality.
- Early Evaluation: Suitable for assessing language coverage, behavior, and overall quality in a preview context.
Important Considerations
This is a preview experiment and is not fully safety-aligned. Users should be aware that the model may produce harmful, biased, or insensitive content. It is not recommended for safety-critical or high-stakes applications without additional safety mitigations. The model is released under an adaptation of the Apache 2.0 license with specific restrictions regarding data recreation and use with licensed press publications.