AmberYifan/capsd-marin-8b-base-code_qurating_b16000_s0

TEXT GENERATIONPricing:Input $0.37 / Cached $0.074 / Output $0.38Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Aug 6, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The AmberYifan/capsd-marin-8b-base-code_qurating_b16000_s0 model is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. It was trained on the capsd_marin-8b-base-n80000-opc__mix_code_qurating_b16000_s0 dataset, suggesting a specialization in code-related tasks. With an 8192 token context length, this model is likely optimized for processing and generating code, or for tasks requiring extensive code understanding.

Loading preview...

Model Overview

The AmberYifan/capsd-marin-8b-base-code_qurating_b16000_s0 is an 8 billion parameter language model, fine-tuned from the marin-community/marin-8b-base architecture. This model has a context length of 8192 tokens, making it suitable for tasks requiring substantial input or output.

Key Training Details

The model underwent a fine-tuning process using the capsd_marin-8b-base-n80000-opc__mix_code_qurating_b16000_s0 dataset. The training utilized specific hyperparameters:

  • Learning Rate: 1e-05
  • Batch Size: A train_batch_size of 2 and eval_batch_size of 8, with a total_train_batch_size of 64 due to gradient accumulation.
  • Optimizer: ADAMW_TORCH with default betas and epsilon.
  • Scheduler: Cosine learning rate scheduler with 0.03 warmup steps.
  • Epochs: Trained for 1 epoch.

Potential Use Cases

Given its fine-tuning on a dataset with "code_qurating" in its name, this model is likely intended for applications involving:

  • Code generation
  • Code completion
  • Code analysis or understanding
  • Refactoring or debugging assistance

Further details on specific capabilities, intended uses, and limitations are not provided in the original model card, suggesting that users should conduct their own evaluations for specific applications.