NbAiLab/nb-gpt-gemma4-26b-a4b-instruct-aurora-2604-pre-aurora-sft-2606-post-epoch3

VISIONConcurrent Unit Cost:2Model Size:26BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 27, 2026Architecture:Transformer Featherless Exclusive Cold

NbAiLab/nb-gpt-gemma4-26b-a4b-instruct-aurora-2604-pre-aurora-sft-2606-post-epoch3 is a 26 billion parameter instruction-tuned language model based on the Gemma 4 architecture, developed by NbAiLab. This model has undergone continued pre-training and subsequent Supervised Fine-Tuning (SFT) on the aurora-sft-2606 dataset, with a context length of 32768 tokens. It is specifically optimized for instruction-following tasks, leveraging assistant-only loss masking with Gemma's tokenizer chat template.

Loading preview...

Model Overview

NbAiLab/nb-gpt-gemma4-26b-a4b-instruct-aurora-2604-pre-aurora-sft-2606-post-epoch3 is a 26 billion parameter instruction-tuned language model. It is built upon the Gemma 4 26B-A4B architecture and has been further developed by NbAiLab through a process of continued pre-training and Supervised Fine-Tuning (SFT).

Key Capabilities

  • Instruction Following: The model is specifically fine-tuned for instruction-following tasks, making it suitable for conversational AI and command-based interactions.
  • Gemma 4 Architecture: Leverages the Gemma 4 base model, indicating a foundation designed for robust language understanding and generation.
  • Extended Context Length: Supports a context length of 32768 tokens, allowing for processing and generating longer texts while maintaining coherence.
  • Optimized Fine-tuning: The SFT process on the NbAiLab/aurora-sft-2606 dataset, combined with assistant-only loss masking using Gemma's chat template, refines its ability to generate relevant and helpful responses in an instructional context.

Good For

  • Instruction-based applications: Ideal for chatbots, virtual assistants, and other systems where the model needs to respond accurately to user instructions.
  • Long-context tasks: Its 32768-token context window makes it suitable for tasks requiring understanding or generation of extensive documents, code, or conversations.
  • Research and Development: Provides a strong base for further fine-tuning or experimentation on instruction-tuned Gemma 4 models.