ferrazzipietro/Llama-3.1-8B-Instruct-reas-int-065-3-epochs-es
The ferrazzipietro/Llama-3.1-8B-Instruct-reas-int-065-3-epochs-es model is a fine-tuned 8 billion parameter instruction-following language model based on Meta's Llama-3.1-8B-Instruct architecture. This iteration was trained for 3 epochs with a cosine learning rate schedule and a total batch size of 64. While specific differentiators and primary use cases are not detailed in the available information, its base model is known for strong general-purpose language understanding and generation.
Loading preview...
Overview
This model, ferrazzipietro/Llama-3.1-8B-Instruct-reas-int-065-3-epochs-es, is a fine-tuned variant of the Meta Llama-3.1-8B-Instruct base model. It leverages the 8 billion parameter architecture of Llama 3.1, which is recognized for its robust performance in various language tasks.
Training Details
The fine-tuning process involved 3 epochs, utilizing a learning rate of 5e-06 and an AdamW optimizer. The training was distributed across 2 GPUs with a total effective batch size of 64, achieved through a gradient_accumulation_steps of 32. A cosine learning rate scheduler with a 0.1 warmup ratio was employed to optimize the training trajectory. The specific dataset used for fine-tuning is not detailed in the provided information.
Key Characteristics
- Base Model: Meta Llama-3.1-8B-Instruct
- Parameters: 8 Billion
- Training Epochs: 3
- Optimizer: AdamW_TORCH with betas=(0.9, 0.95) and epsilon=1e-12
- Learning Rate: 5e-06
- Context Length: 32768 tokens (inherited from base model)
Limitations
Detailed information regarding the model's specific intended uses, limitations, and the dataset used for fine-tuning is not available in the current documentation. Users should exercise caution and conduct further evaluation for specific applications.