msaavedra1234/tiny_t
The msaavedra1234/tiny_t is a 1.1 billion parameter causal language model, fine-tuned from TinyLlama/TinyLlama-1.1B-Chat-v1.0. This model is developed by msaavedra1234 and built using Axolotl, featuring a context length of 2048 tokens. It is optimized for general conversational tasks, leveraging its compact size for efficient deployment. The model demonstrates a final validation loss of 1.8806 after 8 epochs of training.
Loading preview...
Model Overview
The msaavedra1234/tiny_t is a 1.1 billion parameter language model, fine-tuned from the TinyLlama/TinyLlama-1.1B-Chat-v1.0 base model. Developed by msaavedra1234, this model was trained using the Axolotl framework, indicating a focus on efficient and customizable fine-tuning processes. It is designed as a causal language model with a tokenizer based on the Llama architecture.
Key Training Details
- Base Model: TinyLlama/TinyLlama-1.1B-Chat-v1.0
- Training Framework: Axolotl (version 0.3.0)
- Sequence Length: 4096 tokens, with sample packing enabled.
- Optimizer: AdamW with 8-bit quantization (
adamw_bnb_8bit). - Learning Rate: 0.0002, utilizing a cosine scheduler with 10 warmup steps.
- Epochs: Trained for 8 epochs with a micro batch size of 2 and gradient accumulation steps of 4, resulting in a total effective batch size of 8.
- Final Validation Loss: Achieved a validation loss of 1.8806.
Intended Use Cases
This model is suitable for applications requiring a compact yet capable language model, particularly for conversational AI or text generation tasks where resource efficiency is a priority. Its fine-tuned nature suggests potential for specific domain adaptation, though the exact dataset used for fine-tuning is not specified in the provided information. The model's small size makes it ideal for deployment on devices with limited computational resources or for rapid prototyping.