NbAiLab/nb-gpt-gemma4-26b-a4b-instruct-aurora-2604-pre-aurora-sft-2606-post-epoch2
NbAiLab/nb-gpt-gemma4-26b-a4b-instruct-aurora-2604-pre-aurora-sft-2606-post-epoch2 is a 26 billion parameter Gemma 4 model, fine-tuned for two epochs on the NbAiLab/aurora-sft-2606 dataset. This instruction-tuned model utilizes Gemma's tokenizer chat template with assistant-only loss masking. It is designed for conversational AI tasks, building upon a continued-pretraining checkpoint.
Loading preview...
Model Overview
This model, NbAiLab/nb-gpt-gemma4-26b-a4b-instruct-aurora-2604-pre-aurora-sft-2606-post-epoch2, is a 26 billion parameter variant of the Gemma 4 architecture. It has undergone continued pre-training and subsequent instruction-tuning for two epochs on the NbAiLab/aurora-sft-2606 dataset.
Key Characteristics
- Architecture: Based on the Gemma 4 model family.
- Parameter Count: Features 26 billion parameters, offering a balance between capability and computational requirements.
- Context Length: Supports a substantial context window of 32,768 tokens, enabling processing of longer inputs and generating more coherent, extended responses.
- Training Details: Fine-tuned using assistant-only loss masking, leveraging Gemma's native tokenizer chat template. The base checkpoint for continued pre-training was
/cluster/work/projects/nn30001k/versae/gemma4-ffpa-sdpa-bench/checkpoints/aurora-sft-2606-32n-sdpa-packed/step-00002132.
Intended Use Cases
This model is primarily suited for instruction-following tasks and conversational AI applications where a robust understanding of prompts and generation of relevant, coherent text is required. Its instruction-tuned nature makes it effective for chatbots, virtual assistants, and other interactive text generation scenarios.