RefalMachine/ruadapt_qwen2.5_3B_ext_cl100k_bpe_32000_full_lr5e4_2k_bs256

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Oct 11, 2024Architecture:Transformer Featherless Exclusive Cold

RefalMachine/ruadapt_qwen2.5_3B_ext_cl100k_bpe_32000_full_lr5e4_2k_bs256 is a 3.1 billion parameter causal language model, fine-tuned from RefalMachine/ruadapt_qwen2.5_3B_ext_cl100k_bpe_32000_mean_init. This model has a context length of 32768 tokens and was trained with a learning rate of 0.0005 over one epoch. It achieves a validation loss of 2.2974 and an accuracy of 0.5218 on its evaluation set, indicating its performance in language generation tasks.

Loading preview...

Model Overview

RefalMachine/ruadapt_qwen2.5_3B_ext_cl100k_bpe_32000_full_lr5e4_2k_bs256 is a 3.1 billion parameter language model, fine-tuned from the ruadapt_qwen2.5_3B_ext_cl100k_bpe_32000_mean_init base model. It features a substantial context length of 32768 tokens, making it suitable for processing longer sequences of text.

Training Details

The model underwent a single epoch of training using an Adam optimizer with specific beta and epsilon values. Key training hyperparameters included a learning rate of 0.0005, a total batch size of 256, and a cosine learning rate scheduler with 100 warmup steps. The training process utilized 64 devices, indicating a distributed training setup.

Performance Metrics

On its evaluation set, the model achieved a final validation loss of 2.2974 and an accuracy of 0.5218. These metrics provide an indication of its performance in the tasks it was fine-tuned for, though the specific dataset remains unspecified.

Intended Uses & Limitations

Further information regarding the model's intended uses, specific capabilities, and known limitations is not detailed in the provided documentation. Users should exercise caution and conduct their own evaluations for specific applications.