NbAiLab/nb-gpt-gemma4-26b-a4b-instruct-aurora-2604-pre
NbAiLab/nb-gpt-gemma4-26b-a4b-instruct-aurora-2604-pre is a continued-pretraining checkpoint of the Google Gemma 4 26B A4B image-text-to-text conditional generation model, developed by NbAiLab. This 26 billion parameter model, with a 32768 token context length, was trained on approximately 90 billion tokens from the Aurora-2604 text dataset. It is designed as a base model for further evaluation and potential instruction tuning, inheriting characteristics from both the Gemma 4 architecture and its specific training data.
Loading preview...
Model Overview
This model, NbAiLab/nb-gpt-gemma4-26b-a4b-instruct-aurora-2604-pre, is a continued-pretraining checkpoint derived from google/gemma-4-26B-A4B-it. Developed by NbAiLab, it represents the first validated checkpoint from a planned 100 billion token training run, having been trained on approximately 90 billion tokens.
Key Characteristics
- Base Architecture: Inherits from
google/gemma-4-26B-A4B-it, a Gemma 4 26B A4B image-text-to-text conditional generation model. - Training Data: Continued pre-training was performed using the
NbAiLab/aurora-2604-opendataset, specifically processed into fixed-length causal language-modeling blocks. - Training Protocol: Utilized a standard causal language-modeling objective over packed text, without supervised instruction tuning or assistant-only loss masking.
- Checkpoint Details: The exported weights correspond to the first completed checkpoint at approximately 90 billion training tokens, with a sequence length of 8192.
Intended Use and Limitations
This model is provided as a continued-pretrained base checkpoint. It is intended for further evaluation and may require subsequent instruction tuning or alignment to suit specific downstream applications. It inherits limitations from the original Gemma 4 base model and the Aurora-2604 data mixture. Users should note that it is not instruction-tuned and lacks SFT-style user/assistant masking.