aliRafik/Mistral_Nemo_Base_2407_finetuned_16bit

TEXT GENERATIONPricing:Input $0.87 / Cached $0.2 / Output $0.99Concurrent Unit Cost:1Model Size:12BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 10, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The aliRafik/Mistral_Nemo_Base_2407_finetuned_16bit is a 12 billion parameter Mistral-based language model developed by aliRafik. This model was finetuned from unsloth/mistral-nemo-base-2407-bnb-4bit, leveraging Unsloth and Huggingface's TRL library for accelerated training. It offers a 32768 token context length and is optimized for tasks benefiting from efficient finetuning processes.

Loading preview...

Model Overview

The aliRafik/Mistral_Nemo_Base_2407_finetuned_16bit is a 12 billion parameter language model, developed by aliRafik. It is based on the Mistral architecture and was specifically finetuned from the unsloth/mistral-nemo-base-2407-bnb-4bit model. This finetuning process utilized Unsloth and Huggingface's TRL library, which enabled a significantly faster training time, reportedly 2x quicker.

Key Characteristics

  • Architecture: Mistral-based, 12 billion parameters.
  • Training Efficiency: Finetuned using Unsloth and Huggingface's TRL library for accelerated training.
  • Context Length: Supports a substantial context window of 32768 tokens.
  • License: Distributed under the Apache-2.0 license.

Potential Use Cases

This model is suitable for applications requiring a powerful Mistral-based LLM that benefits from efficient finetuning. Its accelerated training process suggests it could be a good candidate for developers looking to quickly adapt a base model to specific downstream tasks or datasets, leveraging the performance benefits of the Mistral architecture combined with optimized training techniques.