lordspline/mergestein
lordspline/mergestein is a 1.5 billion parameter causal language model developed by lordspline. This model is a fine-tuned version of an existing lordspline/mergestein base model, optimized for general conversational tasks. It was trained using Axolotl on a diverse dataset including scientific data, wizard conversations, and ultra-interact data, making it suitable for a broad range of text generation and understanding applications.
Loading preview...
lordspline/mergestein: A Fine-Tuned 1.5B Parameter Model
This model, lordspline/mergestein, is a 1.5 billion parameter causal language model developed by lordspline. It represents a fine-tuned iteration of an existing lordspline/mergestein base model, built using the Axolotl framework (version 0.4.1).
Key Training Details
The model was trained for 1 epoch with a learning rate of 0.0001, utilizing a cosine learning rate scheduler with 10 warmup steps. The training process involved a micro batch size of 1 and gradient accumulation steps of 1. Key datasets used for fine-tuning include:
lordspline/scidata(ShareGPT format)lordspline/wizard_v2_196k_unfiltered(ShareGPT format)lordspline/ultrainteract(ShareGPT format)
Training was conducted with a sequence length of 8192 tokens and employed sample packing. The final validation loss achieved was 1.2069.
Intended Uses
Given its training on diverse conversational and scientific datasets, lordspline/mergestein is generally suitable for a variety of natural language processing tasks, including text generation, question answering, and conversational AI. Its 1.5 billion parameters make it a relatively compact model, potentially offering faster inference compared to larger models while still providing robust language capabilities.