AiAF/KJV-LLM-Pretrained-V1.0
AiAF/KJV-LLM-Pretrained-V1.0 is a 7 billion parameter causal language model developed by AiAF, fine-tuned from mistralai/Mistral-7B-v0.1. This model was specifically pretrained on the AiAF/KJV-LLM-pretraining.jsonl dataset, focusing its knowledge domain. With a context length of 4096 tokens, it is optimized for tasks requiring deep understanding and generation within its specialized training corpus.
Loading preview...
Model Overview
AiAF/KJV-LLM-Pretrained-V1.0 is a 7 billion parameter language model, built upon the mistralai/Mistral-7B-v0.1 architecture. It was developed by AiAF and specifically fine-tuned using the AiAF/KJV-LLM-pretraining.jsonl dataset. This pretraining process aimed to imbue the model with specialized knowledge derived from its unique training data.
Training Details
The model was trained using Axolotl version 0.6.0, with a sequence_len of 8192 and sample_packing enabled. Key training hyperparameters included a learning_rate of 5e-06, num_epochs set to 4, and an adamw_bnb_8bit optimizer. The training process involved 4 gradient_accumulation_steps and a micro_batch_size of 2, resulting in a total_train_batch_size of 8. Flash Attention was utilized to enhance training efficiency. The training concluded with a validation loss of 0.0901.
Potential Use Cases
Given its specialized pretraining, this model is likely best suited for applications requiring deep contextual understanding or generation within the domain of its training dataset. Developers should consider its specific pretraining data when evaluating its suitability for tasks, as its performance will be optimized for that particular knowledge domain.