jiaxie/SpectraLLM_Pretrain_second_stage
Jiaxie's SpectraLLM_Pretrain_second_stage is a 7.6 billion parameter language model, fine-tuned from a LLaMA-Factory base model. It specializes in chemical domain understanding, having been trained on a diverse set of chemistry-related datasets including pubchem_pretrain, nmrbank_pretrain, qm9s_pretrain, hnmr_pretrain, and various mass spectrometry datasets. This model is designed for applications requiring deep knowledge and processing of chemical structures and spectroscopic data, offering a 32768 token context length.
Loading preview...
SpectraLLM_Pretrain_second_stage Overview
This model, developed by jiaxie, is a 7.6 billion parameter language model built upon a LLaMA-Factory base architecture. It represents the second stage of pre-training, specifically fine-tuned for the chemical domain. The model leverages a substantial context length of 32768 tokens, making it suitable for processing extensive chemical data.
Key Training Datasets
The model's specialized capabilities stem from its training on a comprehensive collection of chemistry-specific datasets, including:
pubchem_pretrain: Likely for general chemical compound information.nmrbank_pretrain: Focused on Nuclear Magnetic Resonance (NMR) spectroscopy data.qm9s_pretrain: Pertaining to quantum mechanics and molecular properties.hnmr_pretrain: Specifically for Hydrogen-NMR data.ms_fragment_pos_pretrain,ms_neg_40ev_pretrain,ms_pos_10ev_pretrain,ms_pos_20ev_pretrain: Various mass spectrometry datasets, indicating proficiency in analyzing fragmentation patterns and ion spectra.
Intended Use Cases
Given its specialized training, SpectraLLM_Pretrain_second_stage is particularly well-suited for tasks within chemistry, materials science, and pharmaceutical research. Potential applications include:
- Chemical Information Extraction: Analyzing scientific literature or databases for chemical entities, reactions, and properties.
- Spectroscopy Data Interpretation: Assisting in the analysis and prediction of NMR and Mass Spectrometry data.
- Drug Discovery and Design: Supporting tasks related to molecular property prediction or synthesis planning.
- Materials Science: Understanding and predicting properties of novel materials based on their chemical composition.