lordspline/mergestein

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 10, 2024Architecture:Transformer Featherless Exclusive Cold

lordspline/mergestein is a 1.5 billion parameter causal language model developed by lordspline. This model is a fine-tuned version of an existing lordspline/mergestein base model, optimized for general conversational tasks. It was trained using Axolotl on a diverse dataset including scientific data, wizard conversations, and ultra-interact data, making it suitable for a broad range of text generation and understanding applications.

Loading preview...

lordspline/mergestein: A Fine-Tuned 1.5B Parameter Model

This model, lordspline/mergestein, is a 1.5 billion parameter causal language model developed by lordspline. It represents a fine-tuned iteration of an existing lordspline/mergestein base model, built using the Axolotl framework (version 0.4.1).

Key Training Details

The model was trained for 1 epoch with a learning rate of 0.0001, utilizing a cosine learning rate scheduler with 10 warmup steps. The training process involved a micro batch size of 1 and gradient accumulation steps of 1. Key datasets used for fine-tuning include:

  • lordspline/scidata (ShareGPT format)
  • lordspline/wizard_v2_196k_unfiltered (ShareGPT format)
  • lordspline/ultrainteract (ShareGPT format)

Training was conducted with a sequence length of 8192 tokens and employed sample packing. The final validation loss achieved was 1.2069.

Intended Uses

Given its training on diverse conversational and scientific datasets, lordspline/mergestein is generally suitable for a variety of natural language processing tasks, including text generation, question answering, and conversational AI. Its 1.5 billion parameters make it a relatively compact model, potentially offering faster inference compared to larger models while still providing robust language capabilities.