AmberYifan/capsdnum-marin-8b-base-code_random_b8000_s0
The AmberYifan/capsdnum-marin-8b-base-code_random_b8000_s0 model is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. It was trained on the capsd_marin-8b-base-n80000-opc__mix_code_random_b8000_s0 dataset, suggesting an optimization for code-related tasks. This model is designed for applications requiring a compact yet capable model with a focus on code generation or understanding, leveraging its 8192 token context length.
Loading preview...
Overview
This model, named marin-8b-base_code_random_b8000_s0, is an 8 billion parameter language model developed by AmberYifan. It is a fine-tuned version of the marin-community/marin-8b-base model, specifically adapted using the capsd_marin-8b-base-n80000-opc__mix_code_random_b8000_s0 dataset. The training process involved a learning rate of 1e-05, a total batch size of 64, and utilized a cosine learning rate scheduler over 1 epoch.
Key Training Details
- Base Model: marin-community/marin-8b-base
- Dataset:
capsd_marin-8b-base-n80000-opc__mix_code_random_b8000_s0 - Parameters: 8 billion
- Context Length: 8192 tokens
- Optimizer: ADAMW_TORCH with default betas and epsilon
- Epochs: 1
Intended Use Cases
While specific intended uses and limitations are not detailed in the provided model card, the fine-tuning on a dataset with "code" in its name strongly suggests that this model is optimized for tasks related to code generation, comprehension, or analysis. Developers looking for a specialized 8B model for programming-centric applications might find this model suitable, especially given its substantial context window.