NbAiLab/nb-gpt-gemma4-26b-a4b-instruct-aurora-2604-pre-aurora-sft-2606-post-epoch1
NbAiLab/nb-gpt-gemma4-26b-a4b-instruct-aurora-2604-pre-aurora-sft-2606-post-epoch1 is a 26 billion parameter instruction-tuned language model from NbAiLab, based on the Gemma 4 architecture. It has undergone continued pre-training and subsequent Supervised Fine-Tuning (SFT) specifically on the `NbAiLab/aurora-sft-2606` dataset for one epoch. This model is designed for instruction-following tasks, leveraging its specialized fine-tuning to provide relevant and coherent responses within a 32768 token context window.
Loading preview...
Model Overview
This model, nb-gpt-gemma4-26b-a4b-instruct-aurora-2604-pre-aurora-sft-2606-post-epoch1, is a 26 billion parameter instruction-tuned language model developed by NbAiLab. It is built upon the Gemma 4 architecture and features a substantial context length of 32768 tokens.
Key Characteristics
- Architecture: Based on the Gemma 4 model family.
- Parameter Count: 26 billion parameters, offering a balance between performance and computational requirements.
- Context Length: Supports a large context window of 32768 tokens, enabling processing of extensive inputs and generating detailed outputs.
- Training Regimen: Underwent continued pre-training from a Gemma 4 26B-A4B checkpoint, followed by Supervised Fine-Tuning (SFT) for one epoch on the
NbAiLab/aurora-sft-2606dataset. This specialized fine-tuning aims to enhance its instruction-following capabilities. - Loss Masking: Utilizes assistant-only loss masking, configured with Gemma's tokenizer chat template, to optimize learning for generating helpful assistant responses.
Ideal Use Cases
- Instruction Following: Well-suited for tasks requiring the model to adhere to specific instructions and generate targeted outputs.
- Conversational AI: Its instruction-tuned nature makes it applicable for chatbots and interactive AI applications where coherent and contextually relevant responses are crucial.
- Long-Context Applications: The 32768 token context window makes it suitable for tasks involving extensive documents, detailed queries, or multi-turn conversations.