msaavedra1234/tiny_t

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.1BQuant:BF16Context Size:2kPublished:Jan 4, 2024License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The msaavedra1234/tiny_t is a 1.1 billion parameter causal language model, fine-tuned from TinyLlama/TinyLlama-1.1B-Chat-v1.0. This model is developed by msaavedra1234 and built using Axolotl, featuring a context length of 2048 tokens. It is optimized for general conversational tasks, leveraging its compact size for efficient deployment. The model demonstrates a final validation loss of 1.8806 after 8 epochs of training.

Loading preview...

Model Overview

The msaavedra1234/tiny_t is a 1.1 billion parameter language model, fine-tuned from the TinyLlama/TinyLlama-1.1B-Chat-v1.0 base model. Developed by msaavedra1234, this model was trained using the Axolotl framework, indicating a focus on efficient and customizable fine-tuning processes. It is designed as a causal language model with a tokenizer based on the Llama architecture.

Key Training Details

  • Base Model: TinyLlama/TinyLlama-1.1B-Chat-v1.0
  • Training Framework: Axolotl (version 0.3.0)
  • Sequence Length: 4096 tokens, with sample packing enabled.
  • Optimizer: AdamW with 8-bit quantization (adamw_bnb_8bit).
  • Learning Rate: 0.0002, utilizing a cosine scheduler with 10 warmup steps.
  • Epochs: Trained for 8 epochs with a micro batch size of 2 and gradient accumulation steps of 4, resulting in a total effective batch size of 8.
  • Final Validation Loss: Achieved a validation loss of 1.8806.

Intended Use Cases

This model is suitable for applications requiring a compact yet capable language model, particularly for conversational AI or text generation tasks where resource efficiency is a priority. Its fine-tuned nature suggests potential for specific domain adaptation, though the exact dataset used for fine-tuning is not specified in the provided information. The model's small size makes it ideal for deployment on devices with limited computational resources or for rapid prototyping.