lakshyaixi/Llama_3_2_3B_DPO_v18_refined
lakshyaixi/Llama_3_2_3B_DPO_v18_refined is a 3.2 billion parameter Llama 3 model developed by lakshyaixi, fine-tuned using DPO. This model was trained 2x faster with Unsloth and Huggingface's TRL library, building upon a conversational SFT model. It is designed for refined conversational applications, leveraging its efficient training methodology.
Loading preview...
Model Overview
lakshyaixi/Llama_3_2_3B_DPO_v18_refined is a 3.2 billion parameter Llama 3 model developed by lakshyaixi. It is a fine-tuned variant, specifically utilizing Direct Preference Optimization (DPO) for its refinement. The base model for this DPO finetuning was lakshyaixi/Llama_3_2_3B_Conversational_v6_SFT_10voicebot_interrupt_model, indicating a focus on conversational capabilities.
Key Characteristics
- Architecture: Llama 3 family.
- Parameter Count: 3.2 billion parameters.
- Context Length: Supports a context length of 32768 tokens.
- Training Efficiency: This model was trained significantly faster, achieving a 2x speedup, by leveraging the Unsloth library in conjunction with Huggingface's TRL library.
- Fine-tuning Method: Utilizes DPO (Direct Preference Optimization) for enhanced performance and alignment.
- License: Distributed under the Apache-2.0 license.
Intended Use Cases
This model is well-suited for applications requiring a refined conversational AI, particularly where the base SFT model's conversational and voicebot interrupt capabilities are beneficial. Its efficient training process makes it a practical choice for deployment in scenarios demanding a capable yet relatively compact language model.