nota-ai/cpt-lora_st-vicuna-v1.3-2.7b-ppl
The nota-ai/cpt-lora_st-vicuna-v1.3-2.7b-ppl model, developed by Nota AI, is a 2.7 billion parameter language model derived from Vicuna-v1.3-7B through aggressive depth pruning (60%) and subsequent retraining. It utilizes continued pretraining (CPT) on SlimPajama-627B and further instruction tuning with Low-Rank Adaptation (LoRA) on Refined Alpaca. This model is optimized for efficient text generation by significantly reducing model size while aiming to recover performance, making it suitable for resource-constrained environments.
Loading preview...
Model Overview
This model, cpt-lora_st-vicuna-v1.3-2.7b-ppl, is a 2.7 billion parameter language model developed by Nota AI. It is a depth-pruned version of the Vicuna-v1.3-7B model, specifically reduced by 60% through one-shot pruning based on perplexity (PPL) criterion. The model undergoes a two-stage retraining process: first, continued pretraining (CPT) on the large-scale SlimPajama-627B dataset to recover quality, and then instruction tuning with Low-Rank Adaptation (LoRA) on the Refined Alpaca dataset.
Key Features and Training
- Aggressive Pruning: Achieves a 60% reduction in depth from the original 7B Vicuna-v1.3 model, resulting in a 2.7B parameter model.
- Continued Pretraining (CPT): Retrained on 150 billion tokens from SlimPajama-627B over 12 days using 8 NVIDIA H100 GPUs to mitigate performance degradation from pruning.
- LoRA Instruction Tuning: Further fine-tuned using LoRA on the Refined Alpaca dataset, a process that is highly efficient, requiring only 2 hours and 22GB VRAM on a single NVIDIA A100 GPU for a 20%-pruned model.
- Efficiency Focus: The methodology aims to provide efficient text generation by significantly reducing model size while maintaining performance through targeted retraining.
Intended Use
This model is intended for research and non-commercial projects only, as specified by its license. It is particularly well-suited for applications where computational resources are limited, and a smaller, yet capable, language model is required.