oolabs/sommerfugl-31b

VISIONPricing:Input $0.48 / Cached $0.1 / Output $1.44Concurrent Unit Cost:2Model Size:31BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 30, 2026License:gemmaArchitecture:Transformer Featherless Exclusive Cold

Sommerfugl-31B is a 31 billion parameter Norwegian language model developed by oolabs.no, fine-tuned from Google's Gemma-4-31B-it. This model specifically repairs and enhances Norwegian language knowledge and summarization capabilities, which were degraded in the base Gemma instruction-tuned model. It significantly outperforms the base model and other Norwegian models like Borealis 2 on various Norwegian language tasks, including summarization, grammar correction, and reading comprehension. Sommerfugl-31B is optimized for applications requiring strong performance in Norwegian language understanding and generation.

Loading preview...

Sommerfugl-31B: Enhanced Norwegian Language Model

Sommerfugl-31B, developed by oolabs.no, is a 31 billion parameter language model built upon Google's gemma-4-31B-it. The primary goal of Sommerfugl-31B is to address and significantly improve the Norwegian language capabilities that were observed to regress in the instruction-tuned base Gemma model.

Key Capabilities & Performance

This model demonstrates substantial improvements across critical Norwegian language tasks, as evaluated under the National Library of Norway's leaderboard protocol. It repairs deficiencies in Norwegian language knowledge and summarization, where the base Gemma model showed significant drops in ranking.

  • Superior Norwegian Language Knowledge: Achieves higher accuracy on NoCoLA (0.859) and NCB (0.825) compared to gemma-4-31B-it and Borealis 2.
  • Enhanced Summarization: Shows marked improvements in NorSumm (e.g., 7.39 BLEU for nob) and 'Summarize on request' (4.61 BLEU), significantly outperforming the base model.
  • Improved Grammar and Idiom Handling: Excels in grammar correction (0.402 exact match) and idiom understanding (e.g., 0.442 fscore for nob).
  • Strong Reading Comprehension: Achieves 0.748 F1 on NorQuAD.
  • Overall Outperformance: In head-to-head comparisons across 21 protocol-complete rows, Sommerfugl-31B won 18–3 against Borealis 2 preview.

Known Limitations

  • TruthfulQA (mc) Regression: The model shows a regression in TruthfulQA compared to the base instruction-tuned model, suggesting some erosion of RLHF truthfulness calibration.
  • MMLU-nb Slight Regression: A minor decrease in MMLU Norwegian accuracy (-0.035) compared to the base.

Training data was rigorously quality-gated and decontaminated against all 19 Norwegian evaluation datasets to ensure evaluation integrity and prevent inflated scores due to train/test overlap.