Grogros/dm-llama3.2-1BI-LucieFr-Al4-OWT-TV-ablation-h2d4

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:1Model Size:1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:May 13, 2025License:llama3.2Architecture:Transformer0.0K Featherless Exclusive Warm

Grogros/dm-llama3.2-1BI-LucieFr-Al4-OWT-TV-ablation-h2d4 is a fine-tuned language model based on Meta Llama 3.2-1B-Instruct architecture. This model is a specialized iteration, fine-tuned from the base Llama 3.2-1B-Instruct model. It was trained with specific hyperparameters including a learning rate of 2e-05 and a total batch size of 64 over 2500 training steps. The primary focus of this model is its fine-tuning process, which differentiates it from the original Llama 3.2-1B-Instruct.

Loading preview...

Overview

This model, dm-llama3.2-1BI-LucieFr-Al4-OWT-TV-ablation-h2d4, is a fine-tuned variant of the meta-llama/Llama-3.2-1B-Instruct base model. It leverages the Llama 3.2 architecture, specifically the 1 billion parameter instruction-tuned version, as its foundation. The fine-tuning process involved a specific set of hyperparameters to adapt the model for particular tasks, though the exact dataset used for this fine-tuning is not specified in the available information.

Training Details

The model underwent 2500 training steps with a learning rate of 2e-05. Key training hyperparameters include:

  • Learning Rate: 2e-05
  • Train Batch Size: 4
  • Eval Batch Size: 8
  • Gradient Accumulation Steps: 16
  • Total Train Batch Size: 64
  • Optimizer: ADAFACTOR
  • LR Scheduler Type: cosine with a warmup ratio of 0.1

Intended Use & Limitations

While specific intended uses and limitations are not detailed, as a fine-tuned instruction model, it is generally suitable for tasks requiring instruction following and natural language understanding. Developers should be aware that without further information on the fine-tuning dataset, its performance on specific downstream tasks may vary. Further evaluation is recommended to determine its suitability for particular applications.