AmberYifan/capsd-marin-8b-base-science_random_b80000_s0
The AmberYifan/capsd-marin-8b-base-science_random_b80000_s0 is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. This model is specifically adapted for scientific domains, having been trained on a dataset combining scientific web content and StackExchange data. It is optimized for tasks requiring knowledge and understanding within scientific contexts, leveraging its 8192 token context length.
Loading preview...
Model Overview
This model, AmberYifan/capsd-marin-8b-base-science_random_b80000_s0, is an 8 billion parameter language model. It is a fine-tuned variant of the marin-community/marin-8b-base architecture, specifically adapted for scientific applications.
Key Characteristics
- Base Model: Fine-tuned from
marin-community/marin-8b-base. - Training Data: Specialized fine-tuning on the
capsd_marin-8b-base-n80000-sciweb-stackexchange__mix_science_random_b80000_s0dataset, which includes a mix of scientific web content and StackExchange data. - Parameters: 8 billion parameters.
- Context Length: Supports an 8192 token context window.
Training Details
The model was trained with a learning rate of 1e-05, using a cosine learning rate scheduler with 0.03 warmup steps over 1 epoch. The training utilized a multi-GPU setup with 4 devices, a total batch size of 64, and the AdamW optimizer.
Intended Use Cases
While specific intended uses are not detailed in the original model card, its fine-tuning on scientific datasets suggests suitability for tasks requiring domain-specific knowledge in science, such as:
- Scientific text generation.
- Answering science-related questions.
- Summarization of scientific articles.
- Information extraction from scientific literature.