yoon112/kanana-1.5-8b-instruct-2505-Safe-DPO
The yoon112/kanana-1.5-8b-instruct-2505-Safe-DPO is an 8 billion parameter instruction-tuned Llama model developed by yoon112. This model was fine-tuned using Unsloth and Huggingface's TRL library, resulting in a 2x faster training process. It is designed for general instruction-following tasks, leveraging its efficient training methodology.
Loading preview...
Model Overview
The yoon112/kanana-1.5-8b-instruct-2505-Safe-DPO is an 8 billion parameter Llama-based instruction-tuned model developed by yoon112. This model stands out due to its efficient training process, which was accelerated by 2x using the Unsloth library in conjunction with Huggingface's TRL library. The model is licensed under Apache-2.0.
Key Characteristics
- Architecture: Llama-based, instruction-tuned.
- Parameter Count: 8 billion parameters.
- Context Length: Supports a context length of 8192 tokens.
- Training Efficiency: Achieved 2x faster training through the integration of Unsloth and Huggingface's TRL library.
Use Cases
This model is suitable for a variety of general instruction-following tasks where a balance between performance and computational efficiency during training is desired. Its optimized training process makes it an interesting candidate for developers looking for efficiently produced instruction-tuned models.