AstroMLab/astrollama-3-8b-base_aic

TEXT GENERATIONPricing:Input $0.37 / Cached $0.074 / Output $0.38Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:May 3, 2024License:mitArchitecture:Transformer0.0K Open Weights Featherless Exclusive Cold

AstroMLab's AstroLLaMA-3-8B-Base_AIC is an 8 billion parameter base language model, fine-tuned from Meta's LLaMA-3-8b architecture. Specialized for astronomy, it was trained on Abstract, Introduction, and Conclusion sections from arXiv astro-ph papers. This model is designed for next token prediction in astronomy-related text generation and analysis, rather than instruction following or chat. It offers competitive performance within its class for specialized astronomical tasks.

Loading preview...

AstroLLaMA-3-8B-Base_AIC: Specialized Astronomy Language Model

AstroMLab's AstroLLaMA-3-8B-Base_AIC is an 8 billion parameter base language model built upon Meta's LLaMA-3-8b architecture. It has undergone Continual Pre-Training (CPT) using the LMFlow framework, specifically fine-tuned on astronomical literature. The training data comprises Abstract, Introduction, and Conclusion (AIC) sections extracted from arXiv's astro-ph category papers, processed using OCR and summarization techniques.

Key Characteristics & Use Cases

  • Specialized Domain: Primarily focused on astronomy, making it suitable for generating and analyzing text within this scientific field.
  • Next Token Prediction: Designed for base language modeling tasks, not as an instruction-tuned or chat model.
  • Training Data: Utilizes a unique dataset derived from scientific literature, aiming to imbue it with domain-specific knowledge.
  • Reproducibility: Released to support research reproducibility and allow for tracking the development of AstroLLaMA models.

Performance and Limitations

While AstroLLaMA-3-8B-Base_AIC performs competitively among models of its size on astronomical benchmarking Q&A, it does not surpass the base LLaMA-3-8B model's performance. This indicates that training solely on astro-ph data may not be sufficient for significant gains over highly performant general-purpose models. For optimal performance in astronomy-related tasks, AstroMLab recommends their newer AstroSage-8B model, which addresses these limitations with expanded training data and fine-tuning. Users should verify model outputs against peer-reviewed sources due to potential for generating misleading scientific content.