AmberYifan/capsd-marin-8b-base-math_dsir_b4000_s0

TEXT GENERATIONPricing:Input $0.37 / Cached $0.074 / Output $0.38Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Jul 29, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The AmberYifan/capsd-marin-8b-base-math_dsir_b4000_s0 model is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. It was specifically trained on the capsd_marin-8b-base-n80000-numina__mix_math_dsir_b4000_s0 dataset, suggesting an optimization for mathematical reasoning tasks. This model is intended for applications requiring specialized performance in mathematical domains.

Loading preview...

Model Overview

The AmberYifan/capsd-marin-8b-base-math_dsir_b4000_s0 is an 8 billion parameter language model derived from the marin-community/marin-8b-base architecture. This model has undergone a specific fine-tuning process using the capsd_marin-8b-base-n80000-numina__mix_math_dsir_b4000_s0 dataset.

Training Details

The fine-tuning procedure involved the following hyperparameters:

  • Learning Rate: 1e-05
  • Batch Size: A train_batch_size of 2 and eval_batch_size of 8, leading to a total_train_batch_size of 64 and total_eval_batch_size of 32, utilizing 4 GPUs.
  • Optimizer: ADAMW_TORCH with default betas and epsilon.
  • Scheduler: Cosine learning rate scheduler with 0.03 warmup steps.
  • Epochs: Trained for 1 epoch.

Intended Use

While specific intended uses and limitations are not detailed in the provided information, the dataset name strongly implies that this model is specialized for tasks involving mathematical reasoning and problem-solving. Developers should consider its fine-tuning on a math-centric dataset when evaluating its suitability for their applications.