Grogros/dm-llama3.2-1BI-LucieFr-Al4-OWT-TV-ablation-h2d4
Grogros/dm-llama3.2-1BI-LucieFr-Al4-OWT-TV-ablation-h2d4 is a fine-tuned language model based on Meta Llama 3.2-1B-Instruct architecture. This model is a specialized iteration, fine-tuned from the base Llama 3.2-1B-Instruct model. It was trained with specific hyperparameters including a learning rate of 2e-05 and a total batch size of 64 over 2500 training steps. The primary focus of this model is its fine-tuning process, which differentiates it from the original Llama 3.2-1B-Instruct.
Loading preview...
Overview
This model, dm-llama3.2-1BI-LucieFr-Al4-OWT-TV-ablation-h2d4, is a fine-tuned variant of the meta-llama/Llama-3.2-1B-Instruct base model. It leverages the Llama 3.2 architecture, specifically the 1 billion parameter instruction-tuned version, as its foundation. The fine-tuning process involved a specific set of hyperparameters to adapt the model for particular tasks, though the exact dataset used for this fine-tuning is not specified in the available information.
Training Details
The model underwent 2500 training steps with a learning rate of 2e-05. Key training hyperparameters include:
- Learning Rate: 2e-05
- Train Batch Size: 4
- Eval Batch Size: 8
- Gradient Accumulation Steps: 16
- Total Train Batch Size: 64
- Optimizer: ADAFACTOR
- LR Scheduler Type: cosine with a warmup ratio of 0.1
Intended Use & Limitations
While specific intended uses and limitations are not detailed, as a fine-tuned instruction model, it is generally suitable for tasks requiring instruction following and natural language understanding. Developers should be aware that without further information on the fine-tuning dataset, its performance on specific downstream tasks may vary. Further evaluation is recommended to determine its suitability for particular applications.