AmberYifan/capsd-marin-8b-base-code_random_b14000_s0
AmberYifan/capsd-marin-8b-base-code_random_b14000_s0 is an 8 billion parameter language model fine-tuned from marin-community/marin-8b-base. This model was fine-tuned on the capsd_marin-8b-base-n80000-opc__mix_code_random_b14000_s0 dataset, suggesting an optimization for code-related tasks. It utilizes a context length of 8192 tokens and was trained with a learning rate of 1e-05 over one epoch.
Loading preview...
Model Overview
AmberYifan/capsd-marin-8b-base-code_random_b14000_s0 is an 8 billion parameter language model derived from the marin-community/marin-8b-base architecture. This model has undergone a specific fine-tuning process, utilizing the capsd_marin-8b-base-n80000-opc__mix_code_random_b14000_s0 dataset. While specific details on the dataset's composition are not provided, its naming convention strongly implies a focus on code-related data, suggesting the model is optimized for code generation, completion, or understanding tasks.
Training Details
The fine-tuning process involved a single epoch with a learning rate of 1e-05. Key hyperparameters included a train_batch_size of 2 and an eval_batch_size of 8, with a total effective training batch size of 64 due to gradient accumulation steps. The AdamW optimizer with cosine learning rate scheduling was employed. The model was trained using Transformers 5.7.0 and Pytorch 2.13.0+cu130.
Intended Use
Given its fine-tuning on a code-centric dataset, this model is likely intended for applications requiring strong code understanding and generation capabilities. Developers seeking a specialized 8B parameter model for programming-related tasks may find this model suitable.