donoway/TinyStoriesV2_Llama-3.2-1B-szl5puwo
donoway/TinyStoriesV2_Llama-3.2-1B-szl5puwo is a 1 billion parameter causal language model, fine-tuned from meta-llama/Llama-3.2-1B. This model demonstrates a final validation loss of 1.0261 and a perplexity of 2.7900, with a token accuracy of 0.7067. Its primary characteristic is its compact size, making it suitable for applications requiring a smaller, efficient language model.
Loading preview...
Model Overview
This model, donoway/TinyStoriesV2_Llama-3.2-1B-szl5puwo, is a fine-tuned version of the meta-llama/Llama-3.2-1B architecture, featuring 1 billion parameters. It was trained over 100 epochs with a learning rate of 0.0002 and a batch size of 16.
Key Performance Metrics
During evaluation, the model achieved notable results:
- Final Validation Loss: 1.0261
- Perplexity: 2.7900
- Token Accuracy: 0.7067
- Total Tokens Processed: 2,640,694
Training Details
The training process involved 100 epochs, utilizing the AdamW optimizer with specific beta and epsilon values. The learning rate scheduler was set to a constant type with a warmup ratio of 1e-05. The model's performance metrics steadily improved throughout the training, as evidenced by the decreasing validation loss and increasing token accuracy over 16,500 steps.
Intended Uses & Limitations
While specific intended uses and limitations are not detailed in the original README, its compact 1B parameter size suggests potential for resource-constrained environments or tasks where a smaller, efficient model is advantageous. Further information on its specific applications and constraints would require additional context regarding the unknown dataset it was fine-tuned on.