jiaxie/SpectraLLM_Pretrain_third_stage
The jiaxie/SpectraLLM_Pretrain_third_stage is a 7.6 billion parameter language model, fine-tuned from a second-stage pre-trained model. It was trained on a diverse set of mass spectrometry and nuclear magnetic resonance spectroscopy datasets, including ms_neg_10ev_pretrain, ms_neg_20ev_pretrain, ms_pos_40ev_pretrain, cnmr_pretrain, hsqc_pretrain, ir_pretrain, and ms_fragment_neg_pretrain. This model is specialized for tasks related to chemical spectroscopy data analysis and interpretation, leveraging its 32768 token context length for complex spectral patterns.
Loading preview...
SpectraLLM_Pretrain_third_stage Overview
This model, developed by jiaxie, is a 7.6 billion parameter language model representing the third stage of pre-training within the SpectraLLM series. It builds upon a previously pre-trained second-stage model, further specializing its capabilities through extensive fine-tuning.
Key Capabilities and Training
The model has been fine-tuned on a comprehensive collection of chemical spectroscopy datasets, indicating a strong focus on understanding and processing spectral data. The training involved:
- Diverse Spectroscopy Data: Utilized datasets such as
ms_neg_10ev_pretrain,ms_neg_20ev_pretrain,ms_pos_40ev_pretrain(mass spectrometry),cnmr_pretrain,hsqc_pretrain(NMR spectroscopy),ir_pretrain(infrared spectroscopy), andms_fragment_neg_pretrain. - Training Configuration: Employed a learning rate of 0.0001, a total batch size of 2048 across 32 GPUs with 8 gradient accumulation steps, and a cosine learning rate scheduler with a 0.1 warmup ratio over 1 epoch.
Intended Use Cases
Given its specialized training on various spectroscopy datasets, this model is likely intended for applications requiring advanced understanding and generation related to chemical spectral analysis. Potential use cases include:
- Interpreting mass spectrometry data.
- Analyzing NMR and IR spectroscopy results.
- Assisting in chemical structure elucidation based on spectral patterns.
Further details on specific performance metrics and broader limitations are not provided in the current model card.