lakshyaixi/Llama_3_2_3B_DPO_v18_refined

TEXT GENERATIONConcurrent Unit Cost:1Model Size:3.2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 18, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

lakshyaixi/Llama_3_2_3B_DPO_v18_refined is a 3.2 billion parameter Llama 3 model developed by lakshyaixi, fine-tuned using DPO. This model was trained 2x faster with Unsloth and Huggingface's TRL library, building upon a conversational SFT model. It is designed for refined conversational applications, leveraging its efficient training methodology.

Loading preview...

Model Overview

lakshyaixi/Llama_3_2_3B_DPO_v18_refined is a 3.2 billion parameter Llama 3 model developed by lakshyaixi. It is a fine-tuned variant, specifically utilizing Direct Preference Optimization (DPO) for its refinement. The base model for this DPO finetuning was lakshyaixi/Llama_3_2_3B_Conversational_v6_SFT_10voicebot_interrupt_model, indicating a focus on conversational capabilities.

Key Characteristics

  • Architecture: Llama 3 family.
  • Parameter Count: 3.2 billion parameters.
  • Context Length: Supports a context length of 32768 tokens.
  • Training Efficiency: This model was trained significantly faster, achieving a 2x speedup, by leveraging the Unsloth library in conjunction with Huggingface's TRL library.
  • Fine-tuning Method: Utilizes DPO (Direct Preference Optimization) for enhanced performance and alignment.
  • License: Distributed under the Apache-2.0 license.

Intended Use Cases

This model is well-suited for applications requiring a refined conversational AI, particularly where the base SFT model's conversational and voicebot interrupt capabilities are beneficial. Its efficient training process makes it a practical choice for deployment in scenarios demanding a capable yet relatively compact language model.