AmberYifan/capsd-marin-8b-base-code_kcenter_b12000_s0
The AmberYifan/capsd-marin-8b-base-code_kcenter_b12000_s0 model is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. It was trained on the capsd_marin-8b-base-n80000-opc__mix_code_kcenter_b12000_s0 dataset, suggesting an optimization for code-related tasks. This model utilizes a context length of 8192 tokens and was trained with a learning rate of 1e-05 over one epoch.
Loading preview...
Model Overview
AmberYifan/capsd-marin-8b-base-code_kcenter_b12000_s0 is an 8 billion parameter language model, fine-tuned from the marin-community/marin-8b-base architecture. This model was specifically trained on the capsd_marin-8b-base-n80000-opc__mix_code_kcenter_b12000_s0 dataset, indicating a focus on code-centric applications.
Key Training Details
- Base Model:
marin-community/marin-8b-base - Dataset:
capsd_marin-8b-base-n80000-opc__mix_code_kcenter_b12000_s0 - Parameters: 8 billion
- Context Length: 8192 tokens
- Learning Rate: 1e-05
- Optimizer: ADAMW_TORCH
- Scheduler: Cosine LR scheduler with 0.03 warmup steps
- Epochs: 1
- Batch Size: 2 (train), 8 (eval) with 8 gradient accumulation steps, resulting in a total train batch size of 64.
Intended Use Cases
While specific intended uses and limitations are not detailed in the provided information, the fine-tuning on a code-related dataset suggests its potential suitability for tasks such as:
- Code generation
- Code completion
- Code summarization
- Debugging assistance
Users should be aware that detailed performance metrics and further specific use cases are not yet available and would require additional evaluation.