AmberYifan/capsd-marin-8b-base-science_random_b2000_s0

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Jul 16, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The AmberYifan/capsd-marin-8b-base-science_random_b2000_s0 model is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. This model was specifically trained on a scientific dataset, capsd_marin-8b-base-n10000__mix_science_random_b2000_s0, with a context length of 8192 tokens. It is optimized for tasks related to scientific text processing and understanding, leveraging its base architecture and specialized fine-tuning.

Loading preview...

Model Overview

The AmberYifan/capsd-marin-8b-base-science_random_b2000_s0 is an 8 billion parameter language model, fine-tuned from the marin-community/marin-8b-base architecture. This model has been specialized through training on the capsd_marin-8b-base-n10000__mix_science_random_b2000_s0 dataset, indicating a focus on scientific domains.

Training Details

The fine-tuning process involved specific hyperparameters:

  • Learning Rate: 1e-05
  • Batch Sizes: train_batch_size of 1, eval_batch_size of 8, with a gradient_accumulation_steps of 16, resulting in a total_train_batch_size of 64.
  • Optimizer: ADAMW_TORCH with default betas and epsilon.
  • Scheduler: Cosine learning rate scheduler with 0.03 warmup steps.
  • Epochs: Trained for 1 epoch.

Intended Use

Given its fine-tuning on a scientific dataset, this model is likely best suited for applications involving scientific text, such as information extraction from research papers, scientific question answering, or generating scientific summaries. Its 8B parameters and 8192 token context length provide a solid foundation for processing complex scientific information.