RefalMachine/mistral_darulm_20_05_24_part1-2_32000_unigram_part1_lr1e4_bs256
RefalMachine/mistral_darulm_20_05_24_part1-2_32000_unigram_part1_lr1e4_bs256 is a 7 billion parameter language model, fine-tuned from RefalMachine/mistral_darulm_20_05_24_part1-2_32000_unigram_mean_init_03_07_24. This model was trained for 1 epoch with a learning rate of 0.0001 and achieved a validation loss of 2.0924 and an accuracy of 0.5528 on an unspecified dataset. Its primary characteristic is its fine-tuned nature, though specific differentiators or intended uses are not detailed in the provided information.
Loading preview...
Model Overview
RefalMachine/mistral_darulm_20_05_24_part1-2_32000_unigram_part1_lr1e4_bs256 is a 7 billion parameter language model, fine-tuned from an existing RefalMachine model. The fine-tuning process involved 1 epoch of training with a learning rate of 0.0001, utilizing Adam optimizer with specific beta and epsilon values, and a cosine learning rate scheduler with 100 warmup steps. The training was distributed across 32 devices, resulting in a total batch size of 128 for both training and evaluation.
Training Performance
During its single epoch of training, the model achieved a final validation loss of 2.0924 and an accuracy of 0.5528. The training progress showed a consistent decrease in validation loss and an increase in accuracy over 22,000 steps.
Key Training Hyperparameters
- Learning Rate: 0.0001
- Optimizer: Adam (betas=(0.9, 0.95), epsilon=1e-05)
- Scheduler: Cosine with 100 warmup steps
- Epochs: 1.0
- Total Batch Size: 128 (across 32 devices)
Limitations
The model card indicates that more information is needed regarding its specific intended uses, limitations, and the details of the training and evaluation datasets. Therefore, its optimal application areas and potential biases are currently undefined.