lakshyaixi/Llama_3_2_3B_DPO_v21_on_v20_0622prod
The lakshyaixi/Llama_3_2_3B_DPO_v21_on_v20_0622prod is a 3.2 billion parameter Llama 3 model, developed by lakshyaixi, with a 32768 token context length. This model was finetuned using Unsloth and Huggingface's TRL library, achieving 2x faster training. It is a DPO-finetuned iteration, building upon a previous version, and is suitable for applications requiring a compact yet capable Llama 3 variant.
Loading preview...
Model Overview
The lakshyaixi/Llama_3_2_3B_DPO_v21_on_v20_0622prod is a 3.2 billion parameter Llama 3 model developed by lakshyaixi. This iteration, version v21, is a DPO (Direct Preference Optimization) finetune based on the lakshyaixi/Llama_3_2_3B_DPO_v20_on_v7_loopcleaned model.
Key Characteristics
- Architecture: Llama 3 base model.
- Parameter Count: 3.2 billion parameters, offering a balance between performance and computational efficiency.
- Context Length: Supports a substantial context window of 32768 tokens.
- Training Efficiency: Finetuned using Unsloth and Huggingface's TRL library, which enabled 2x faster training compared to standard methods.
- Finetuning Method: Utilizes Direct Preference Optimization (DPO) for alignment.
Intended Use Cases
This model is suitable for applications where a compact, DPO-aligned Llama 3 variant is beneficial. Its optimized training process suggests potential for efficient deployment in various NLP tasks, particularly those benefiting from preference-based finetuning.