lakshyaixi/Llama_3_2_3B_DPO_antiloop_v2
TEXT GENERATIONConcurrent Unit Cost:1Model Size:3.2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 26, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold
lakshyaixi/Llama_3_2_3B_DPO_antiloop_v2 is a 3.2 billion parameter Llama 3 model developed by lakshyaixi, fine-tuned using DPO. This model was trained 2x faster with Unsloth and Huggingface's TRL library, building upon lakshyaixi/Llama_3_2_3B_Conversational_v7_SFT_10voicebot_loopcleaned. It is designed for conversational applications, leveraging its efficient training methodology.
Loading preview...
Model Overview
lakshyaixi/Llama_3_2_3B_DPO_antiloop_v2 is a 3.2 billion parameter Llama 3 model developed by lakshyaixi. This model has been fine-tuned using Direct Preference Optimization (DPO) and builds upon the previously trained lakshyaixi/Llama_3_2_3B_Conversational_v7_SFT_10voicebot_loopcleaned.
Key Training Details
- Efficient Training: The model was trained significantly faster, achieving a 2x speedup, by utilizing Unsloth and Huggingface's TRL library. This highlights an optimization in the training process for Llama 3 models.
- Base Model: It is finetuned from a conversational Llama 3 variant, suggesting its primary application area.
Potential Use Cases
- Conversational AI: Given its base model and DPO fine-tuning, it is likely well-suited for dialogue systems, chatbots, and interactive AI applications.
- Research into Efficient Fine-tuning: Developers interested in applying Unsloth for faster Llama 3 model training can examine this model as an example.