oolabs/sommerfugl-31b-v1

VISIONPricing:Input $0.48 / Cached $0.1 / Output $1.44Concurrent Unit Cost:2Model Size:31BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 30, 2026License:gemmaArchitecture:Transformer0.0K Featherless Exclusive Cold

Sommerfugl-31B-v1 by oolabs.no is a 31 billion parameter Norwegian language model, fine-tuned from Google's Gemma-4-31B-it with a 32K context length. It significantly improves Norwegian language knowledge and summarization capabilities, which were degraded in the base Gemma instruction-tuned model. This model is optimized for tasks requiring strong Norwegian linguistic understanding and generation, such as summarization, grammar correction, and idiom comprehension.

Loading preview...

Sommerfugl-31B-v1: Enhanced Norwegian Language Model

Sommerfugl-31B-v1, developed by oolabs.no, is a 31 billion parameter model built upon Google's gemma-4-31B-it. Its primary purpose is to restore and significantly enhance Norwegian language capabilities that were diminished in the instruction-tuned Gemma base model. While gemma-4-31B-it performed poorly on Norwegian language knowledge and summarization tasks, Sommerfugl-31B-v1 repairs these deficiencies.

Key Capabilities & Performance

This model demonstrates substantial improvements across various Norwegian-specific benchmarks:

  • Norwegian Language Knowledge: Achieves 0.859 on NoCoLA and 0.825 on NCB, outperforming the base Gemma and Borealis 2.
  • Summarization: Shows significant gains with 7.39/7.00 BLEU for NorSumm (nob/nno) and 4.61 BLEU for summarization on request.
  • Idiom Comprehension: Improves f-score to 0.442/0.541 (nob/nno).
  • Grammar Correction: Reaches 0.402 exact match accuracy.
  • Reading Comprehension: Scores 0.748 F1 on NorQuAD.
  • Translation: Slightly improves English-Norwegian translation with 59.1 BLEU.

Sommerfugl-31B-v1 notably outperforms NbAiLab's Borealis 2 preview in 18 out of 21 head-to-head benchmark rows. Training data was rigorously decontaminated against all 19 Norwegian evaluation datasets to ensure evaluation integrity.

Known Limitations

  • TruthfulQA (mc): Shows a regression of -0.19 compared to the base model, indicating some erosion of truthfulness calibration.
  • MMLU Norwegian: A slight decrease of -0.035 compared to the base gemma-4-31B-it.

Good for

  • Applications requiring high-quality Norwegian text generation and understanding.
  • Tasks such as summarization, content creation, and grammar correction in Norwegian.
  • Developers seeking a powerful Norwegian-centric LLM with strong benchmark performance.