AstroMLab/astrollama-3-8b-base_aic
AstroMLab's AstroLLaMA-3-8B-Base_AIC is an 8 billion parameter base language model, fine-tuned from Meta's LLaMA-3-8b architecture. Specialized for astronomy, it was trained on Abstract, Introduction, and Conclusion sections from arXiv astro-ph papers. This model is designed for next token prediction in astronomy-related text generation and analysis, rather than instruction following or chat. It offers competitive performance within its class for specialized astronomical tasks.
Loading preview...
AstroLLaMA-3-8B-Base_AIC: Specialized Astronomy Language Model
AstroMLab's AstroLLaMA-3-8B-Base_AIC is an 8 billion parameter base language model built upon Meta's LLaMA-3-8b architecture. It has undergone Continual Pre-Training (CPT) using the LMFlow framework, specifically fine-tuned on astronomical literature. The training data comprises Abstract, Introduction, and Conclusion (AIC) sections extracted from arXiv's astro-ph category papers, processed using OCR and summarization techniques.
Key Characteristics & Use Cases
- Specialized Domain: Primarily focused on astronomy, making it suitable for generating and analyzing text within this scientific field.
- Next Token Prediction: Designed for base language modeling tasks, not as an instruction-tuned or chat model.
- Training Data: Utilizes a unique dataset derived from scientific literature, aiming to imbue it with domain-specific knowledge.
- Reproducibility: Released to support research reproducibility and allow for tracking the development of AstroLLaMA models.
Performance and Limitations
While AstroLLaMA-3-8B-Base_AIC performs competitively among models of its size on astronomical benchmarking Q&A, it does not surpass the base LLaMA-3-8B model's performance. This indicates that training solely on astro-ph data may not be sufficient for significant gains over highly performant general-purpose models. For optimal performance in astronomy-related tasks, AstroMLab recommends their newer AstroSage-8B model, which addresses these limitations with expanded training data and fine-tuning. Users should verify model outputs against peer-reviewed sources due to potential for generating misleading scientific content.