RefalMachine/ruadapt_qwen2.5_3B_ext_cl100k_bpe_32000_full_lr3e4_2k_bs256
RefalMachine/ruadapt_qwen2.5_3B_ext_cl100k_bpe_32000_full_lr3e4_2k_bs256 is a 3.1 billion parameter language model, fine-tuned from RefalMachine's ruadapt_qwen2.5_3B_ext_cl100k_bpe_32000_mean_init. This model, based on the Qwen2.5 architecture, features an extended context length of 32768 tokens. It was trained for 1 epoch with a learning rate of 0.0003, achieving a validation loss of 2.3223 and an accuracy of 0.5190 on its evaluation set.
Loading preview...
Model Overview
RefalMachine/ruadapt_qwen2.5_3B_ext_cl100k_bpe_32000_full_lr3e4_2k_bs256 is a 3.1 billion parameter language model, fine-tuned from the base model RefalMachine/ruadapt_qwen2.5_3B_ext_cl100k_bpe_32000_mean_init. This model leverages the Qwen2.5 architecture and is configured with an extended context length of 32768 tokens, making it suitable for processing longer sequences of text.
Training Details
The model underwent a single epoch of training using a learning rate of 0.0003 and a total batch size of 256 across 64 devices. The training process utilized an Adam optimizer with cosine learning rate scheduling and 100 warmup steps. During evaluation, the model achieved a final validation loss of 2.3223 and an accuracy of 0.5190.
Performance Metrics
Key training results include:
- Final Validation Loss: 2.3223
- Final Accuracy: 0.5190
Framework Versions
The model was trained using:
- Transformers 4.37.2
- Pytorch 2.3.0a0+6ddf5cf85e.nv24.04
- Datasets 2.18.0
- Tokenizers 0.15.2