lakshyaixi/Llama_3_2_3B_DPO_v21_on_v20_0622prod

TEXT GENERATIONConcurrent Unit Cost:1Model Size:3.2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 26, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The lakshyaixi/Llama_3_2_3B_DPO_v21_on_v20_0622prod is a 3.2 billion parameter Llama 3 model, developed by lakshyaixi, with a 32768 token context length. This model was finetuned using Unsloth and Huggingface's TRL library, achieving 2x faster training. It is a DPO-finetuned iteration, building upon a previous version, and is suitable for applications requiring a compact yet capable Llama 3 variant.

Loading preview...

Model Overview

The lakshyaixi/Llama_3_2_3B_DPO_v21_on_v20_0622prod is a 3.2 billion parameter Llama 3 model developed by lakshyaixi. This iteration, version v21, is a DPO (Direct Preference Optimization) finetune based on the lakshyaixi/Llama_3_2_3B_DPO_v20_on_v7_loopcleaned model.

Key Characteristics

  • Architecture: Llama 3 base model.
  • Parameter Count: 3.2 billion parameters, offering a balance between performance and computational efficiency.
  • Context Length: Supports a substantial context window of 32768 tokens.
  • Training Efficiency: Finetuned using Unsloth and Huggingface's TRL library, which enabled 2x faster training compared to standard methods.
  • Finetuning Method: Utilizes Direct Preference Optimization (DPO) for alignment.

Intended Use Cases

This model is suitable for applications where a compact, DPO-aligned Llama 3 variant is beneficial. Its optimized training process suggests potential for efficient deployment in various NLP tasks, particularly those benefiting from preference-based finetuning.