AmberYifan/capsd-convfinqa-fullscore-marin-8b-base-finance_ppl_b4000_s0
The AmberYifan/capsd-convfinqa-fullscore-marin-8b-base-finance_ppl_b4000_s0 model is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. This model is specifically adapted for financial applications, having been trained on the capsd_marin-8b-base-n11082-finance-convfinqa-fullscore__mix_finance_ppl_b4000_s0 dataset. It is designed for tasks requiring financial domain understanding, leveraging its 8192 token context length for processing relevant information.
Loading preview...
Model Overview
This model, AmberYifan/capsd-convfinqa-fullscore-marin-8b-base-finance_ppl_b4000_s0, is an 8 billion parameter language model. It is a fine-tuned variant of the marin-community/marin-8b-base architecture, specifically adapted for financial domain tasks.
Key Characteristics
- Base Model: Fine-tuned from
marin-community/marin-8b-base. - Parameter Count: 8 billion parameters.
- Context Length: Supports an 8192 token context window.
- Domain Specialization: The model has undergone specialized training on a financial dataset,
capsd_marin-8b-base-n11082-finance-convfinqa-fullscore__mix_finance_ppl_b4000_s0, indicating its focus on financial applications.
Training Details
The model was trained with a learning rate of 1e-05, a total batch size of 64 (achieved with a train_batch_size of 2 and gradient_accumulation_steps of 8 across 4 GPUs), and utilized the AdamW optimizer. The training consisted of 1 epoch with a cosine learning rate scheduler and 0.03 warmup steps. It was developed using Transformers 5.7.0 and PyTorch 2.13.0+cu130.
Intended Use
Given its fine-tuning on a financial dataset, this model is intended for use cases requiring deep understanding and generation within the financial sector.