ferrazzipietro/Llama-3.1-8B-Instruct-reas-int-065-3-epochs-it

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 4, 2026License:llama3.1Architecture:Transformer Featherless Exclusive Cold

The ferrazzipietro/Llama-3.1-8B-Instruct-reas-int-065-3-epochs-it model is an 8 billion parameter instruction-tuned language model, fine-tuned from Meta's Llama-3.1-8B-Instruct. This model was trained for 3 epochs with a learning rate of 5e-06 and a context length of 32768 tokens. Its specific differentiation and primary use case are not detailed in the provided information, as it is fine-tuned on an unknown dataset.

Loading preview...

Model Overview

This model, ferrazzipietro/Llama-3.1-8B-Instruct-reas-int-065-3-epochs-it, is an 8 billion parameter instruction-tuned language model. It is a fine-tuned version of the meta-llama/Llama-3.1-8B-Instruct base model.

Training Details

The model underwent a fine-tuning process for 3 epochs using the following key hyperparameters:

  • Learning Rate: 5e-06
  • Batch Size: A train_batch_size of 1 and gradient_accumulation_steps of 32 resulted in a total_train_batch_size of 64.
  • Optimizer: ADAMW_TORCH with betas=(0.9, 0.95) and epsilon=1e-12.
  • Scheduler: Cosine learning rate scheduler with a warmup ratio of 0.1.
  • Distributed Training: Utilized a multi-GPU setup with 2 devices.

Limitations

The specific dataset used for fine-tuning is currently unknown, and further details regarding the model's intended uses, limitations, and evaluation data are not provided in the available documentation. Users should exercise caution and conduct their own evaluations to determine suitability for specific applications.