AmberYifan/capsd-marin-8b-base-science_random_b2000_s0
The AmberYifan/capsd-marin-8b-base-science_random_b2000_s0 model is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. This model was specifically trained on a scientific dataset, capsd_marin-8b-base-n10000__mix_science_random_b2000_s0, with a context length of 8192 tokens. It is optimized for tasks related to scientific text processing and understanding, leveraging its base architecture and specialized fine-tuning.
Loading preview...
Model Overview
The AmberYifan/capsd-marin-8b-base-science_random_b2000_s0 is an 8 billion parameter language model, fine-tuned from the marin-community/marin-8b-base architecture. This model has been specialized through training on the capsd_marin-8b-base-n10000__mix_science_random_b2000_s0 dataset, indicating a focus on scientific domains.
Training Details
The fine-tuning process involved specific hyperparameters:
- Learning Rate: 1e-05
- Batch Sizes:
train_batch_sizeof 1,eval_batch_sizeof 8, with agradient_accumulation_stepsof 16, resulting in atotal_train_batch_sizeof 64. - Optimizer: ADAMW_TORCH with default betas and epsilon.
- Scheduler: Cosine learning rate scheduler with 0.03 warmup steps.
- Epochs: Trained for 1 epoch.
Intended Use
Given its fine-tuning on a scientific dataset, this model is likely best suited for applications involving scientific text, such as information extraction from research papers, scientific question answering, or generating scientific summaries. Its 8B parameters and 8192 token context length provide a solid foundation for processing complex scientific information.