AmberYifan/capsd-opc-dedup-marin-8b-base-code_ppl_b60000_s0
AmberYifan/capsd-opc-dedup-marin-8b-base-code_ppl_b60000_s0 is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. This model has a context length of 8192 tokens and was trained on a specific dataset, capsd_marin-8b-base-n80000-opc-dedup80k__mix_code_ppl_b60000_s0. It is designed for general language understanding and generation tasks, with its fine-tuning potentially enhancing its performance in specific domains related to its training data.
Loading preview...
This model, AmberYifan/capsd-opc-dedup-marin-8b-base-code_ppl_b60000_s0, is an 8 billion parameter language model. It is a fine-tuned variant of the marin-community/marin-8b-base model, specifically trained on the capsd_marin-8b-base-n80000-opc-dedup80k__mix_code_ppl_b60000_s0 dataset.
Overview
This model was fine-tuned using a learning rate of 1e-05, a total training batch size of 64, and a cosine learning rate scheduler with 0.03 warmup steps over 1 epoch. The training utilized a multi-GPU setup with 4 devices.
Key Training Details
- Base Model:
marin-community/marin-8b-base - Dataset:
capsd_marin-8b-base-n80000-opc-dedup80k__mix_code_ppl_b60000_s0 - Parameters: 8 Billion
- Context Length: 8192 tokens
- Optimizer: AdamW with betas=(0.9, 0.999) and epsilon=1e-08
- Epochs: 1
Intended Use
Given its fine-tuning on a specific dataset, this model is likely optimized for tasks related to the characteristics of the capsd_marin-8b-base-n80000-opc-dedup80k__mix_code_ppl_b60000_s0 dataset. Developers should consider its training data for specific applications, as its general capabilities are derived from the base marin-8b-base model.