AmberYifan/capsd-opc-dedup-marin-8b-base-code_random_b1000_s0
AmberYifan/capsd-opc-dedup-marin-8b-base-code_random_b1000_s0 is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. This model was trained on a specialized dataset, capsd_marin-8b-base-n80000-opc-dedup80k__mix_code_random_b1000_s0, with a context length of 8192 tokens. It is optimized for tasks related to code generation and understanding, leveraging its base architecture and specific training data.
Loading preview...
Model Overview
This model, marin-8b-base_code_random_b1000_s0, is an 8 billion parameter language model developed by AmberYifan. It is a fine-tuned variant of the marin-community/marin-8b-base architecture, specifically adapted through further training.
Key Training Details
The model underwent fine-tuning on the capsd_marin-8b-base-n80000-opc-dedup80k__mix_code_random_b1000_s0 dataset. Training was conducted using a learning rate of 1e-05, a total batch size of 64 (with a train batch size of 2 and gradient accumulation steps of 8), and the AdamW optimizer. A cosine learning rate scheduler was employed over 1 epoch, with a warmup ratio of 0.03. The training utilized a multi-GPU setup with 4 devices.
Potential Use Cases
Given its fine-tuning on a code-centric dataset, this model is likely suitable for applications involving:
- Code generation: Assisting in writing or completing code snippets.
- Code understanding: Analyzing and interpreting programming language constructs.
- Code-related tasks: Any task benefiting from a model exposed to a significant volume of code data.