AmberYifan/capsd-opc-dedup-marin-8b-base-code_cap_b2000_s0
AmberYifan/capsd-opc-dedup-marin-8b-base-code_cap_b2000_s0 is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. This model is specifically adapted for code-related tasks, leveraging a deduplicated dataset for its training. It is designed to enhance performance in code generation and understanding within its 8192 token context window. The model's specialization makes it suitable for applications requiring robust code processing capabilities.
Loading preview...
Model Overview
This model, AmberYifan/capsd-opc-dedup-marin-8b-base-code_cap_b2000_s0, is an 8 billion parameter language model derived from the marin-community/marin-8b-base architecture. It has been fine-tuned on a specialized dataset, capsd_marin-8b-base-n80000-opc-dedup80k__mix_code_cap_b2000_s0, indicating an optimization for code-related tasks.
Key Characteristics
- Base Model: Fine-tuned from
marin-community/marin-8b-base. - Parameter Count: 8 billion parameters.
- Context Length: Supports an 8192 token context window.
- Training Data: Utilizes a deduplicated dataset, suggesting a focus on data quality for its specific domain.
Training Details
The model was trained with a learning rate of 1e-05, a train_batch_size of 2, and a gradient_accumulation_steps of 8, resulting in a total_train_batch_size of 64. It used the AdamW_TORCH optimizer and a cosine learning rate scheduler with 0.03 warmup steps over 1 epoch. The training was conducted on a multi-GPU setup with 4 devices.
Good For
- Code Generation: Its fine-tuning on a code-centric dataset suggests proficiency in generating programming code.
- Code Understanding: Likely capable of interpreting and analyzing code snippets effectively.
- Applications requiring specialized code processing: Suitable for tasks where a model with dedicated code training is beneficial.