NotoriousH2/kanana-1.5-8b-instruct-2505-Safe-DPO
The NotoriousH2/kanana-1.5-8b-instruct-2505-Safe-DPO is an 8 billion parameter Llama-based instruction-tuned model developed by NotoriousH2. This model was fine-tuned using Unsloth and Huggingface's TRL library, emphasizing efficient training. It is designed for general instruction-following tasks, leveraging its optimized training process for performance.
Loading preview...
Model Overview
The NotoriousH2/kanana-1.5-8b-instruct-2505-Safe-DPO is an 8 billion parameter instruction-tuned model, developed by NotoriousH2. It is based on the Llama architecture and was fine-tuned from an existing NotoriousH2 model. The training process utilized Unsloth and Huggingface's TRL library, which enabled a 2x faster fine-tuning compared to standard methods.
Key Characteristics
- Architecture: Llama-based, 8 billion parameters.
- Training Efficiency: Fine-tuned with Unsloth, resulting in significantly faster training times.
- Context Length: Supports an 8192-token context window.
- License: Released under the Apache-2.0 license.
Use Cases
This model is suitable for a variety of instruction-following applications where an efficiently trained and performant 8B parameter model is beneficial. Its optimized training process suggests it could be a good choice for developers looking for a capable model with a focus on resource efficiency during fine-tuning.