jordanpainter/diallm-llama-cpt
The jordanpainter/diallm-llama-cpt is an 8 billion parameter Llama 3.1-based model, continually pretrained on 18 varieties of the International Corpus of English (approximately 20 million tokens). Developed by Jordan Painter et al. as part of the DiaLLM project, this checkpoint serves as a foundational model for English dialect adaptation. It utilizes GaLore for memory-efficient full-parameter optimization during continual pretraining, preceding dialect-specific fine-tuning. This model is specifically designed to investigate the robustness-generation gap in English dialect adaptation.
Loading preview...
DiaLLM Llama 3.1-8B Continual Pretraining Checkpoint
This model, jordanpainter/diallm-llama-cpt, is an 8 billion parameter Llama 3.1 base model that has undergone continual pretraining (CPT). It is a core component of the DiaLLM project, which focuses on English dialect adaptation in large language models.
Key Capabilities & Characteristics
- Dialect Adaptation Foundation: Serves as the shared foundational checkpoint before branching into implicit and explicit dialect-specific adaptations.
- Continual Pretraining Data: Trained on the International Corpus of English, encompassing 18 distinct English varieties and approximately 20 million tokens.
- Memory-Efficient Optimization: Leverages GaLore for memory-efficient full-parameter optimization during its continual pretraining phase.
- Research Focus: Developed as part of an investigation into the robustness-generation gap in English dialect adaptation, detailed in the paper "DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation" (arXiv:2607.07669).
Good For
- Researchers and developers working on English dialect adaptation and understanding linguistic variations.
- As a base model for further fine-tuning on specific English dialects.
- Exploring the impact of continual pretraining on diverse linguistic corpora.