ahans1/standard-loss-llama-1B
The ahans1/standard-loss-llama-1B is a 1.1 billion parameter causal language model developed by Abhimanyu Hans and collaborators, based on the TinyLLaMA-1.1B architecture. It was pretrained on 20 billion Redpajama tokens using standard causal language modeling loss. This model serves as a control baseline for research into 'goldfish loss,' a method designed to mitigate memorization in generative LLMs by pseudorandomly dropping tokens during loss computation. Its primary use is as a comparative benchmark for evaluating the effectiveness of novel loss functions in reducing training data memorization.
Loading preview...
Overview of ahans1/standard-loss-llama-1B
The ahans1/standard-loss-llama-1B model is a 1.1 billion parameter language model built upon the TinyLLaMA-1.1B architecture. Developed by Abhimanyu Hans and collaborators, this model was pretrained on 20 billion tokens from the Redpajama dataset using a standard causal language modeling loss function. It serves as a crucial baseline for research detailed in the paper "Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs" (arXiv:2406.10209).
Key Characteristics
- Architecture: Based on the TinyLLaMA-1.1B framework.
- Training Data: Pretrained on 20 billion Redpajama tokens.
- Loss Function: Utilizes standard causal language modeling loss, without the token-dropping mechanism of 'goldfish loss'.
- Context Length: Supports a context length of 2048 tokens.
- Memorization Baseline: Includes a 'canaries dataset' of 2000 Wikipedia documents, repeated 50 times during pre-training, to establish a baseline for memorization behavior under standard training conditions.
Good for
- Research on Memorization: Ideal for comparing against models trained with novel loss functions, such as 'goldfish loss', to evaluate their effectiveness in mitigating training data memorization.
- Understanding Standard LLM Behavior: Provides a representative example of a 1.1B parameter model trained with conventional methods, useful for studying general language model characteristics.
- Benchmarking: Can be used as a control model in experiments assessing the impact of different training methodologies on model performance and data retention.