donoway/TinyStoriesV2_Llama-3.2-1B-szl5puwo

Hugging Face
TEXT GENERATIONPricing:Input $0.108 / Output $0.804Concurrent Unit Cost:1Model Size:1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 17, 2025License:llama3.2Architecture:Transformer Featherless Exclusive Warm

donoway/TinyStoriesV2_Llama-3.2-1B-szl5puwo is a 1 billion parameter causal language model, fine-tuned from meta-llama/Llama-3.2-1B. This model demonstrates a final validation loss of 1.0261 and a perplexity of 2.7900, with a token accuracy of 0.7067. Its primary characteristic is its compact size, making it suitable for applications requiring a smaller, efficient language model.

Loading preview...

Model Overview

This model, donoway/TinyStoriesV2_Llama-3.2-1B-szl5puwo, is a fine-tuned version of the meta-llama/Llama-3.2-1B architecture, featuring 1 billion parameters. It was trained over 100 epochs with a learning rate of 0.0002 and a batch size of 16.

Key Performance Metrics

During evaluation, the model achieved notable results:

  • Final Validation Loss: 1.0261
  • Perplexity: 2.7900
  • Token Accuracy: 0.7067
  • Total Tokens Processed: 2,640,694

Training Details

The training process involved 100 epochs, utilizing the AdamW optimizer with specific beta and epsilon values. The learning rate scheduler was set to a constant type with a warmup ratio of 1e-05. The model's performance metrics steadily improved throughout the training, as evidenced by the decreasing validation loss and increasing token accuracy over 16,500 steps.

Intended Uses & Limitations

While specific intended uses and limitations are not detailed in the original README, its compact 1B parameter size suggests potential for resource-constrained environments or tasks where a smaller, efficient model is advantageous. Further information on its specific applications and constraints would require additional context regarding the unknown dataset it was fine-tuned on.