RefalMachine/mistral_darulm_20_05_24_part1-2_32000_bpe_part1-2_lr1e4_bs256
RefalMachine/mistral_darulm_20_05_24_part1-2_32000_bpe_part1-2_lr1e4_bs256 is a 7 billion parameter language model fine-tuned from RefalMachine/mistral_darulm_20_05_24_part1-2_32000_bpe_mean_init_03_07_24. This model was trained with a 4096 token context length and achieved a final validation loss of 2.0469 and an accuracy of 0.5649. Its specific primary use case and unique differentiators are not detailed in the provided information.
Loading preview...
Model Overview
RefalMachine/mistral_darulm_20_05_24_part1-2_32000_bpe_part1-2_lr1e4_bs256 is a 7 billion parameter language model, fine-tuned from an existing base model, RefalMachine/mistral_darulm_20_05_24_part1-2_32000_bpe_mean_init_03_07_24. The model was trained for 1 epoch with a learning rate of 0.0001 and a total batch size of 256 across 64 devices.
Training Performance
During its single epoch of training, the model demonstrated a consistent reduction in validation loss, ultimately achieving a final validation loss of 2.0469 and an accuracy of 0.5649. The training utilized an Adam optimizer with specific beta and epsilon values, and a cosine learning rate scheduler with 100 warmup steps.
Key Technical Details
- Base Model: Fine-tuned from RefalMachine/mistral_darulm_20_05_24_part1-2_32000_bpe_mean_init_03_07_24.
- Parameters: 7 billion.
- Context Length: 4096 tokens.
- Training Frameworks: Transformers 4.37.2, Pytorch 2.3.0a0+6ddf5cf85e.nv24.04, Datasets 2.18.0, Tokenizers 0.15.2.
Use Cases
Specific intended uses and limitations are not detailed in the provided model card. Developers should evaluate its performance on their specific tasks, considering the reported loss and accuracy metrics, as the training dataset is also unspecified.