ferrazzipietro/Llama-3.2-1B-Instruct-reas-int-065-3-epochs-es
ferrazzipietro/Llama-3.2-1B-Instruct-reas-int-065-3-epochs-es is a 1 billion parameter instruction-tuned language model, fine-tuned by ferrazzipietro from the Meta Llama-3.2-1B-Instruct base model. This model has a context length of 32768 tokens and was trained for 3 epochs with specific hyperparameters including a learning rate of 5e-06. Its primary differentiation and intended use cases are not explicitly detailed in the provided information, suggesting it may be an experimental or specialized fine-tune.
Loading preview...
Model Overview
This model, ferrazzipietro/Llama-3.2-1B-Instruct-reas-int-065-3-epochs-es, is a fine-tuned variant of the Meta Llama-3.2-1B-Instruct base model. It features 1 billion parameters and supports a context length of 32768 tokens. The fine-tuning process involved 3 epochs with a learning rate of 5e-06, a train_batch_size of 4, and gradient_accumulation_steps of 8, resulting in a total_train_batch_size of 64. The optimizer used was ADAMW_TORCH with specific beta and epsilon values, and a cosine learning rate scheduler with a warmup ratio of 0.1.
Training Details
- Base Model: meta-llama/Llama-3.2-1B-Instruct
- Epochs: 3
- Learning Rate: 5e-06
- Optimizer: ADAMW_TORCH (betas=(0.9, 0.95), epsilon=1e-12)
- Scheduler: Cosine with 0.1 warmup ratio
- Frameworks: Transformers 4.57.0, Pytorch 2.14.0+cu130, Datasets 5.0.1, Tokenizers 0.22.2
Current Limitations
The specific dataset used for fine-tuning, the model's intended uses, and its performance characteristics are not detailed in the available information. Users should be aware that without further documentation, its specialized capabilities or optimal use cases are unknown.