AmberYifan/capsd-marin-8b-base-code_random_b12000_s0

TEXT GENERATIONPricing:Input $0.37 / Cached $0.074 / Output $0.38Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Aug 5, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

AmberYifan/capsd-marin-8b-base-code_random_b12000_s0 is an 8 billion parameter language model fine-tuned from marin-community/marin-8b-base. This model was trained on the capsd_marin-8b-base-n80000-opc__mix_code_random_b12000_s0 dataset, indicating a specialization in code-related tasks. It utilizes a cosine learning rate scheduler and was trained for one epoch with a learning rate of 1e-05. This model is intended for applications requiring code generation or understanding, leveraging its base architecture and specific fine-tuning.

Loading preview...

Model Overview

AmberYifan/capsd-marin-8b-base-code_random_b12000_s0 is an 8 billion parameter language model that has been fine-tuned from the marin-community/marin-8b-base architecture. The fine-tuning process specifically utilized the capsd_marin-8b-base-n80000-opc__mix_code_random_b12000_s0 dataset, suggesting an optimization for tasks involving code.

Training Details

The model was trained using the following key hyperparameters:

  • Learning Rate: 1e-05
  • Batch Size: A train_batch_size of 2 and eval_batch_size of 8, leading to a total_train_batch_size of 64 and total_eval_batch_size of 32, with 8 gradient accumulation steps.
  • Optimizer: ADAMW_TORCH with default betas and epsilon.
  • LR Scheduler: Cosine type with 0.03 warmup steps.
  • Epochs: Trained for 1 epoch.

This configuration indicates a focused training approach aimed at enhancing the model's capabilities on the specific fine-tuning dataset. The model was developed using Transformers 5.7.0, Pytorch 2.13.0+cu130, Datasets 4.0.0, and Tokenizers 0.22.2.