AmberYifan/capsd-marin-8b-base-math_less_b4000_s0

TEXT GENERATIONPricing:Input $0.37 / Cached $0.074 / Output $0.38Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Jul 30, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The AmberYifan/capsd-marin-8b-base-math_less_b4000_s0 model is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. This model was trained on the capsd_marin-8b-base-n80000-numina__mix_math_less_b4000_s0 dataset, suggesting a specialization in mathematical reasoning or related tasks. It leverages a cosine learning rate scheduler and AdamW optimizer, indicating a focus on robust and efficient training for its specific domain.

Loading preview...

Model Overview

The AmberYifan/capsd-marin-8b-base-math_less_b4000_s0 is an 8 billion parameter language model, fine-tuned from the marin-community/marin-8b-base architecture. This model was specifically trained on the capsd_marin-8b-base-n80000-numina__mix_math_less_b4000_s0 dataset, which implies a targeted optimization for tasks involving mathematical reasoning or quantitative analysis.

Training Details

The fine-tuning process utilized the following key hyperparameters:

  • Learning Rate: 1e-05
  • Optimizer: AdamW with betas=(0.9, 0.999) and epsilon=1e-08
  • Batch Size: A total training batch size of 64 (2 per device across 4 GPUs with 8 gradient accumulation steps)
  • Scheduler: Cosine learning rate scheduler with 0.03 warmup steps
  • Epochs: Trained for 1 epoch

This configuration suggests a focused and efficient fine-tuning approach to adapt the base model to its specialized dataset.

Intended Use Cases

While specific intended uses and limitations are not detailed in the provided README, the dataset name strongly indicates that this model is likely optimized for:

  • Mathematical problem-solving
  • Quantitative reasoning tasks
  • Applications requiring numerical understanding