5456es/last_layer_prune_Llama-3.1-8B-Instruct_prune_0.7-sigmoid
The 5456es/last_layer_prune_Llama-3.1-8B-Instruct_prune_0.7-sigmoid is an 8 billion parameter instruction-tuned causal language model, fine-tuned from Llama-3.1-8B-Instruct using Direct Preference Optimization (DPO) with a 'last' method pruning strategy. This model is optimized for generating responses based on preference data, inheriting the base model's 32768 token context length. Its primary differentiator is the application of pruning during DPO training, aiming for efficient performance in instruction-following tasks.
Loading preview...
Overview
This model, 5456es/last_layer_prune_Llama-3.1-8B-Instruct_prune_0.7-sigmoid, is an 8 billion parameter instruction-tuned language model. It is a fine-tuned variant of the Llama-3.1-8B-Instruct base model, specifically optimized using the Direct Preference Optimization (DPO) algorithm. A key characteristic of this model is the application of a 'last' method pruning strategy during its DPO training, which differentiates it from standard DPO fine-tunes.
Key Capabilities
- Instruction Following: Inherits and refines the instruction-following capabilities of its Llama-3.1-8B-Instruct base.
- Preference Alignment: Fine-tuned on preference data using DPO, aiming to generate responses aligned with human preferences.
- Pruning Integration: Incorporates a pruning process during training, potentially leading to more efficient inference while maintaining performance.
Good For
- Research into Pruning and DPO: Ideal for researchers exploring the effects of pruning techniques combined with DPO for model optimization.
- Instruction-tuned Applications: Suitable for applications requiring a model that can follow instructions effectively, particularly where efficiency gains from pruning are beneficial.
- Comparative Studies: Can be used to compare the performance and efficiency of DPO models with and without integrated pruning strategies.