AnkitBirGurung/NEMO-12B-SFT-Further

TEXT GENERATIONConcurrent Unit Cost:1Model Size:12BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 2, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

AnkitBirGurung/NEMO-12B-SFT-Further is a 12 billion parameter Mistral-based model, fine-tuned from Roxwell/NEMO-12B-SFT. Developed by Roxwell, this model was trained using Unsloth and Huggingface's TRL library, achieving 2x faster training speeds. With a 32768 token context length, it is optimized for efficient performance in language generation tasks.

Loading preview...

Model Overview

AnkitBirGurung/NEMO-12B-SFT-Further is a 12 billion parameter language model, fine-tuned from the Roxwell/NEMO-12B-SFT base model. Developed by Roxwell, this model leverages the Mistral architecture and features a substantial context length of 32768 tokens.

Key Characteristics

  • Efficient Training: This model was trained significantly faster (2x) by utilizing Unsloth and Huggingface's TRL library, indicating an optimization for training efficiency.
  • Base Model: Fine-tuned from Roxwell/NEMO-12B-SFT, suggesting a focus on specific instruction-following or conversational capabilities inherited from its predecessor.
  • Context Length: Supports a 32768 token context window, enabling the processing of longer inputs and maintaining coherence over extended conversations or documents.

Good For

  • Applications requiring a 12B parameter model with a large context window.
  • Scenarios where efficient training methodologies are a point of interest or a performance indicator.
  • Tasks benefiting from a Mistral-based architecture that has undergone further supervised fine-tuning.