ljh728/kanana-1.5-8b-instruct-2505-Safe-DPO

TEXT GENERATIONPricing:Input $0.37 / Cached $0.074 / Output $0.38Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Sep 4, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The ljh728/kanana-1.5-8b-instruct-2505-Safe-DPO is an 8 billion parameter Llama-based instruction-tuned language model developed by ljh728. This model was fine-tuned using Unsloth and Huggingface's TRL library, enabling a 2x faster training process. It is designed for general instruction-following tasks, leveraging its efficient training methodology.

Loading preview...

Model Overview

The ljh728/kanana-1.5-8b-instruct-2505-Safe-DPO is an 8 billion parameter instruction-tuned language model based on the Llama architecture. Developed by ljh728, this model stands out due to its efficient training methodology, utilizing Unsloth and Huggingface's TRL library. This combination allowed for a 2x faster finetuning process compared to standard methods.

Key Characteristics

  • Architecture: Llama-based, instruction-tuned.
  • Parameter Count: 8 billion parameters.
  • Training Efficiency: Leverages Unsloth for significantly faster finetuning.
  • License: Distributed under the Apache-2.0 license.

Good For

  • General Instruction Following: Capable of handling a wide range of instruction-based prompts.
  • Applications requiring efficient models: Benefits from its optimized training, potentially leading to more streamlined deployment.
  • Developers interested in Unsloth: Serves as an example of a model finetuned with Unsloth for improved speed.