AmberYifan/capsd-opc-dedup-marin-8b-base-code_random_b8000_s0
AmberYifan/capsd-opc-dedup-marin-8b-base-code_random_b8000_s0 is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. This model was specifically trained on the capsd_marin-8b-base-n80000-opc-dedup80k__mix_code_random_b8000_s0 dataset, suggesting an optimization for code-related tasks. It utilizes an 8192 token context length and was trained with a learning rate of 1e-05 over one epoch.
Loading preview...
Model Overview
AmberYifan/capsd-opc-dedup-marin-8b-base-code_random_b8000_s0 is an 8 billion parameter language model, fine-tuned from the marin-community/marin-8b-base architecture. This model was specifically adapted using the capsd_marin-8b-base-n80000-opc-dedup80k__mix_code_random_b8000_s0 dataset.
Training Details
The model underwent a fine-tuning process with the following key hyperparameters:
- Learning Rate: 1e-05
- Batch Size: A total training batch size of 64 (2 per device across 4 GPUs with 8 gradient accumulation steps).
- Optimizer: ADAMW_TORCH with default betas and epsilon.
- Scheduler: Cosine learning rate scheduler with 0.03 warmup steps.
- Epochs: Trained for 1 epoch.
Potential Use Cases
Given its fine-tuning on a dataset with "code_random" in its name, this model is likely intended for applications involving code generation, completion, or analysis. Its 8B parameter size and 8192 token context length make it suitable for tasks requiring moderate complexity and context understanding in programming domains.