RefalMachine/ruadapt_qwen2.5_3B_ext_cl100k_bpe_32000_full_lr3e4_2k_bs256

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Oct 11, 2024Architecture:Transformer Featherless Exclusive Cold

RefalMachine/ruadapt_qwen2.5_3B_ext_cl100k_bpe_32000_full_lr3e4_2k_bs256 is a 3.1 billion parameter language model, fine-tuned from RefalMachine's ruadapt_qwen2.5_3B_ext_cl100k_bpe_32000_mean_init. This model, based on the Qwen2.5 architecture, features an extended context length of 32768 tokens. It was trained for 1 epoch with a learning rate of 0.0003, achieving a validation loss of 2.3223 and an accuracy of 0.5190 on its evaluation set.

Loading preview...

Model Overview

RefalMachine/ruadapt_qwen2.5_3B_ext_cl100k_bpe_32000_full_lr3e4_2k_bs256 is a 3.1 billion parameter language model, fine-tuned from the base model RefalMachine/ruadapt_qwen2.5_3B_ext_cl100k_bpe_32000_mean_init. This model leverages the Qwen2.5 architecture and is configured with an extended context length of 32768 tokens, making it suitable for processing longer sequences of text.

Training Details

The model underwent a single epoch of training using a learning rate of 0.0003 and a total batch size of 256 across 64 devices. The training process utilized an Adam optimizer with cosine learning rate scheduling and 100 warmup steps. During evaluation, the model achieved a final validation loss of 2.3223 and an accuracy of 0.5190.

Performance Metrics

Key training results include:

  • Final Validation Loss: 2.3223
  • Final Accuracy: 0.5190

Framework Versions

The model was trained using:

  • Transformers 4.37.2
  • Pytorch 2.3.0a0+6ddf5cf85e.nv24.04
  • Datasets 2.18.0
  • Tokenizers 0.15.2