AmberYifan/capsd-marin-8b-base-code_kcenter_b10000_s0
The AmberYifan/capsd-marin-8b-base-code_kcenter_b10000_s0 is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. This model is specifically optimized for code-related tasks, having been trained on the capsd_marin-8b-base-n80000-opc__mix_code_kcenter_b10000_s0 dataset. It features an 8192-token context length and is designed for applications requiring code generation or understanding.
Loading preview...
Model Overview
The AmberYifan/capsd-marin-8b-base-code_kcenter_b10000_s0 is an 8 billion parameter language model, fine-tuned from the marin-community/marin-8b-base architecture. This model has been specialized through fine-tuning on the capsd_marin-8b-base-n80000-opc__mix_code_kcenter_b10000_s0 dataset, indicating an optimization for code-centric applications.
Key Training Details
The model underwent a single epoch of training with a learning rate of 1e-05. Key hyperparameters include:
- Optimizer: ADAMW_TORCH with betas=(0.9, 0.999) and epsilon=1e-08
- Batch Size: A total training batch size of 64 (train_batch_size: 2, gradient_accumulation_steps: 8)
- LR Scheduler: Cosine type with 0.03 warmup steps
- Distributed Training: Utilized 4 GPUs
Intended Uses
Given its fine-tuning on a code-specific dataset, this model is likely intended for tasks such as:
- Code generation
- Code completion
- Code understanding and analysis
Limitations
The model card indicates that more information is needed regarding its specific limitations and detailed intended uses. Users should perform thorough evaluations for their specific use cases.