ferrazzipietro/Llama-3.2-1B-Instruct-cpt-tesi_all
The ferrazzipietro/Llama-3.2-1B-Instruct-cpt-tesi_all is a 1 billion parameter instruction-tuned causal language model, fine-tuned by ferrazzipietro from the meta-llama/Llama-3.2-1B-Instruct base model. This model has a context length of 32768 tokens. It is designed for general instruction-following tasks, leveraging its Llama-3.2 architecture for efficient processing.
Loading preview...
Model Overview
This model, Llama-3.2-1B-Instruct-cpt-tesi_all, is a 1 billion parameter instruction-tuned language model developed by ferrazzipietro. It is a fine-tuned variant of the meta-llama/Llama-3.2-1B-Instruct base model, designed to follow instructions effectively. The model was trained with a learning rate of 0.0002, a total batch size of 1024, and for 1 epoch, utilizing a multi-GPU setup.
Key Training Details
- Base Model:
meta-llama/Llama-3.2-1B-Instruct - Parameters: 1 billion
- Context Length: 32768 tokens
- Learning Rate: 0.0002
- Optimizer: ADAFACTOR
- Scheduler: Cosine with 0.3 warmup ratio
- Epochs: 1
- Frameworks: Transformers 4.57.0, Pytorch 2.4.0+cu118, Datasets 3.6.0, Tokenizers 0.22.2
Intended Use Cases
While specific intended uses and limitations are not detailed in the provided information, as an instruction-tuned model, it is generally suitable for a variety of natural language processing tasks that involve following explicit instructions. Its 1 billion parameters make it a relatively compact model, potentially offering faster inference compared to larger models, which could be beneficial for resource-constrained environments or applications requiring quick responses.