NbAiLab/nb-gpt-gemma4-26b-a4b-instruct-aurora-2604-pre

VISIONConcurrent Unit Cost:2Model Size:26BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 27, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

NbAiLab/nb-gpt-gemma4-26b-a4b-instruct-aurora-2604-pre is a continued-pretraining checkpoint of the Google Gemma 4 26B A4B image-text-to-text conditional generation model, developed by NbAiLab. This 26 billion parameter model, with a 32768 token context length, was trained on approximately 90 billion tokens from the Aurora-2604 text dataset. It is designed as a base model for further evaluation and potential instruction tuning, inheriting characteristics from both the Gemma 4 architecture and its specific training data.

Loading preview...

Model Overview

This model, NbAiLab/nb-gpt-gemma4-26b-a4b-instruct-aurora-2604-pre, is a continued-pretraining checkpoint derived from google/gemma-4-26B-A4B-it. Developed by NbAiLab, it represents the first validated checkpoint from a planned 100 billion token training run, having been trained on approximately 90 billion tokens.

Key Characteristics

  • Base Architecture: Inherits from google/gemma-4-26B-A4B-it, a Gemma 4 26B A4B image-text-to-text conditional generation model.
  • Training Data: Continued pre-training was performed using the NbAiLab/aurora-2604-open dataset, specifically processed into fixed-length causal language-modeling blocks.
  • Training Protocol: Utilized a standard causal language-modeling objective over packed text, without supervised instruction tuning or assistant-only loss masking.
  • Checkpoint Details: The exported weights correspond to the first completed checkpoint at approximately 90 billion training tokens, with a sequence length of 8192.

Intended Use and Limitations

This model is provided as a continued-pretrained base checkpoint. It is intended for further evaluation and may require subsequent instruction tuning or alignment to suit specific downstream applications. It inherits limitations from the original Gemma 4 base model and the Aurora-2604 data mixture. Users should note that it is not instruction-tuned and lacks SFT-style user/assistant masking.