AmberYifan/capsd-marin-8b-base-science_ppl_b2000_s0
AmberYifan/capsd-marin-8b-base-science_ppl_b2000_s0 is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. This model was specifically trained on a science-related dataset, indicating an optimization for scientific text processing and understanding. It utilizes a context length of 8192 tokens, making it suitable for tasks requiring analysis of moderately long scientific documents.
Loading preview...
Model Overview
AmberYifan/capsd-marin-8b-base-science_ppl_b2000_s0 is an 8 billion parameter language model derived from the marin-community/marin-8b-base architecture. This model has undergone a specific fine-tuning process using the capsd_marin-8b-base-n10000__mix_science_ppl_b2000_s0 dataset, suggesting a specialization in scientific domains.
Training Details
The fine-tuning was conducted with the following key hyperparameters:
- Learning Rate: 1e-05
- Batch Size: A total training batch size of 64 (achieved with
train_batch_size: 1andgradient_accumulation_steps: 16across 4 devices). - Optimizer: ADAMW_TORCH with default betas and epsilon.
- LR Scheduler: Cosine scheduler with 0.03 warmup steps.
- Epochs: Trained for 1 epoch.
Potential Use Cases
Given its fine-tuning on a science-specific dataset, this model is likely optimized for tasks involving:
- Processing and generating scientific text.
- Understanding scientific concepts and terminology.
- Applications within scientific research or education where domain-specific language comprehension is crucial.