AmberYifan/capsd-less-ultra-humaneval-opc-marin-8b-base-code_less_b1000_s0
AmberYifan/capsd-less-ultra-humaneval-opc-marin-8b-base-code_less_b1000_s0 is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. This model was trained on the capsd_marin-8b-base-n80000-opc__mix_code_less_b1000_s0 dataset, suggesting a specialization in code-related tasks. With an 8192 token context length, it is designed for applications requiring processing of moderately long code sequences.
Loading preview...
Model Overview
AmberYifan/capsd-less-ultra-humaneval-opc-marin-8b-base-code_less_b1000_s0 is an 8 billion parameter language model, fine-tuned from the marin-community/marin-8b-base architecture. It was specifically trained on the capsd_marin-8b-base-n80000-opc__mix_code_less_b1000_s0 dataset.
Training Details
The model underwent a single epoch of training with a learning rate of 1e-05. Key hyperparameters included a train_batch_size of 2, an eval_batch_size of 8, and a gradient_accumulation_steps of 8, resulting in a total_train_batch_size of 64. The optimizer used was ADAMW_TORCH with a cosine learning rate scheduler and 0.03 warmup steps. The training utilized 4 GPUs.
Current Status
As of its release, further details regarding the model's specific description, intended uses, limitations, and comprehensive training and evaluation data are pending. Users are encouraged to consult future updates for more in-depth information on its capabilities and optimal applications.