AmberYifan/capsdnum-marin-8b-base-code_cap_b8000_s0

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Jul 16, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The AmberYifan/capsdnum-marin-8b-base-code_cap_b8000_s0 model is a fine-tuned version of the marin-community/marin-8b-base architecture. It was specifically trained on the capsd_marin-8b-base-n80000-opc__mix_code_cap_b8000_s0 dataset, indicating a specialization in code-related tasks. This model is optimized for code generation and understanding, leveraging its base model's capabilities with further refinement on a dedicated code dataset. It is suitable for applications requiring robust performance in programming contexts.

Loading preview...

Model Overview

This model, AmberYifan/capsdnum-marin-8b-base-code_cap_b8000_s0, is a specialized fine-tuned variant of the marin-community/marin-8b-base model. Its development focused on enhancing performance for code-related applications through targeted training.

Training Details

The model was fine-tuned using the capsd_marin-8b-base-n80000-opc__mix_code_cap_b8000_s0 dataset. Key training hyperparameters included:

  • Learning Rate: 1e-05
  • Batch Size: 2 (train), 8 (eval)
  • Gradient Accumulation Steps: 8
  • Optimizer: AdamW with betas=(0.9, 0.999) and epsilon=1e-08
  • LR Scheduler: Cosine type with 0.03 warmup steps
  • Epochs: 1

The training was conducted across 4 multi-GPU devices, utilizing Transformers 5.7.0, Pytorch 2.13.0+cu130, Datasets 4.0.0, and Tokenizers 0.22.2.

Intended Use

Given its fine-tuning on a code-specific dataset, this model is primarily intended for tasks involving code generation, completion, analysis, or other programming-centric applications. Further details on specific use cases and limitations are not explicitly provided in the current model card.