RefalMachine/ruadapt_qwen2.5_3B_ext_cl100k_bpe_32000_full_lr2e4_2k_bs256

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Oct 11, 2024Architecture:Transformer Featherless Exclusive Cold

RefalMachine/ruadapt_qwen2.5_3B_ext_cl100k_bpe_32000_full_lr2e4_2k_bs256 is a 3.1 billion parameter language model, fine-tuned from RefalMachine's ruadapt_qwen2.5_3B_ext_cl100k_bpe_32000_mean_init. This model has a context length of 32768 tokens and was trained with a learning rate of 0.0002 over one epoch. It achieved a final validation loss of 2.3590 and an accuracy of 0.5146, indicating its performance on the specific fine-tuning task.

Loading preview...

Overview

This model, ruadapt_qwen2.5_3B_ext_cl100k_bpe_32000_full_lr2e4_2k_bs256, is a fine-tuned variant of RefalMachine's ruadapt_qwen2.5_3B_ext_cl100k_bpe_32000_mean_init model. It is a 3.1 billion parameter language model designed with a substantial context length of 32768 tokens, making it suitable for processing longer sequences of text.

Training Details

The model underwent a single epoch of fine-tuning using a learning rate of 0.0002 and a total batch size of 256 across 64 devices. The training process utilized an Adam optimizer with cosine learning rate scheduling and 100 warmup steps. Over 22,000 steps, the model's validation loss improved from 5.2631 to 2.3591, with accuracy increasing from 0.3345 to 0.5146.

Performance Metrics

  • Final Validation Loss: 2.3590
  • Final Accuracy: 0.5146

Frameworks Used

  • Transformers: 4.37.2
  • Pytorch: 2.3.0a0+6ddf5cf85e.nv24.04
  • Datasets: 2.18.0
  • Tokenizers: 0.15.2