mbakshi1094/llama-2-7b-logit-watermark-distill-kgw-k1-gamma0.25-delta2-hk22983996
The mbakshi1094/llama-2-7b-logit-watermark-distill-kgw-k1-gamma0.25-delta2-hk22983996 is a 7 billion parameter Llama-2-7b-hf model, fine-tuned by mbakshi1094 on the Skylion007/openwebtext dataset. This model is designed for general text generation tasks, leveraging the Llama 2 architecture. Its training on a broad web text corpus suggests applicability for diverse language understanding and generation applications.
Loading preview...
Model Overview
This model, llama-2-7b-logit-watermark-distill-kgw-k1-gamma0.25-delta2-hk22983996, is a fine-tuned variant of the meta-llama/Llama-2-7b-hf base model. It features 7 billion parameters and was trained by mbakshi1094.
Training Details
The model was fine-tuned on the Skylion007/openwebtext dataset. Key training hyperparameters include:
- Learning Rate: 1e-05
- Batch Size: 16 (train), 8 (eval)
- Optimizer: AdamW_Torch with betas=(0.9, 0.999) and epsilon=1e-08
- Scheduler: Cosine learning rate scheduler with 500 warmup steps
- Training Steps: 5000
- Frameworks: Transformers 4.57.6, Pytorch 2.4.1+cu124, Datasets 4.8.5, Tokenizers 0.22.2
Potential Use Cases
Given its foundation on Llama 2 and fine-tuning on a general web text dataset, this model is suitable for a range of natural language processing tasks, including:
- Text generation
- Language understanding
- Content creation