AmberYifan/capsd-less-humaneval-opc-marin-8b-base-code_less_b1000_s0
AmberYifan/capsd-less-humaneval-opc-marin-8b-base-code_less_b1000_s0 is an 8 billion parameter language model fine-tuned from marin-community/marin-8b-base. This model was trained on the capsd_marin-8b-base-n80000-opc__mix_code_less_b1000_s0 dataset, suggesting a specialization in code-related tasks. It utilizes a context length of 8192 tokens and was fine-tuned for one epoch with a learning rate of 1e-05.
Loading preview...
Model Overview
AmberYifan/capsd-less-humaneval-opc-marin-8b-base-code_less_b1000_s0 is an 8 billion parameter language model, fine-tuned from the marin-community/marin-8b-base architecture. This model was specifically trained on the capsd_marin-8b-base-n80000-opc__mix_code_less_b1000_s0 dataset, indicating a focus on code-related applications.
Key Training Details
- Base Model: Fine-tuned from
marin-community/marin-8b-base. - Dataset: Trained on
capsd_marin-8b-base-n80000-opc__mix_code_less_b1000_s0. - Parameters: 8 billion.
- Context Length: 8192 tokens.
- Training Epochs: 1 epoch.
- Learning Rate: 1e-05.
- Optimizer: ADAMW_TORCH with betas=(0.9, 0.999) and epsilon=1e-08.
- Batch Size: A total training batch size of 64 (with 2 per device and 8 gradient accumulation steps across 4 GPUs).
Potential Use Cases
Given its fine-tuning on a code-centric dataset, this model is likely intended for tasks such as:
- Code generation.
- Code completion.
- Code understanding and analysis.
Limitations
The model card indicates that more information is needed regarding its intended uses and limitations, as well as detailed training and evaluation data. Users should perform their own evaluations to determine suitability for specific applications.