ahans1/standard-loss-llama-1B

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.1BQuant:BF16Context Size:2kPublished:May 8, 2024License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The ahans1/standard-loss-llama-1B is a 1.1 billion parameter causal language model developed by Abhimanyu Hans and collaborators, based on the TinyLLaMA-1.1B architecture. It was pretrained on 20 billion Redpajama tokens using standard causal language modeling loss. This model serves as a control baseline for research into 'goldfish loss,' a method designed to mitigate memorization in generative LLMs by pseudorandomly dropping tokens during loss computation. Its primary use is as a comparative benchmark for evaluating the effectiveness of novel loss functions in reducing training data memorization.

Loading preview...

Overview of ahans1/standard-loss-llama-1B

The ahans1/standard-loss-llama-1B model is a 1.1 billion parameter language model built upon the TinyLLaMA-1.1B architecture. Developed by Abhimanyu Hans and collaborators, this model was pretrained on 20 billion tokens from the Redpajama dataset using a standard causal language modeling loss function. It serves as a crucial baseline for research detailed in the paper "Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs" (arXiv:2406.10209).

Key Characteristics

  • Architecture: Based on the TinyLLaMA-1.1B framework.
  • Training Data: Pretrained on 20 billion Redpajama tokens.
  • Loss Function: Utilizes standard causal language modeling loss, without the token-dropping mechanism of 'goldfish loss'.
  • Context Length: Supports a context length of 2048 tokens.
  • Memorization Baseline: Includes a 'canaries dataset' of 2000 Wikipedia documents, repeated 50 times during pre-training, to establish a baseline for memorization behavior under standard training conditions.

Good for

  • Research on Memorization: Ideal for comparing against models trained with novel loss functions, such as 'goldfish loss', to evaluate their effectiveness in mitigating training data memorization.
  • Understanding Standard LLM Behavior: Provides a representative example of a 1.1B parameter model trained with conventional methods, useful for studying general language model characteristics.
  • Benchmarking: Can be used as a control model in experiments assessing the impact of different training methodologies on model performance and data retention.