sanchit-gandhi/distil-zephyr-1.5b-ssft-ultrachat
The sanchit-gandhi/distil-zephyr-1.5b-ssft-ultrachat model is a fine-tuned version of sanchit-gandhi/Mistral-7B-v0.1-6-layer, developed by sanchit-gandhi. This 7 billion parameter model was fine-tuned on the Ultrachat dataset. It is designed for general conversational AI tasks, leveraging its training on a diverse chat corpus.
Loading preview...
Overview
This model, developed by sanchit-gandhi, is a fine-tuned variant of the sanchit-gandhi/Mistral-7B-v0.1-6-layer base model. It has been specifically adapted through supervised fine-tuning (SSFT) using the stingning/ultrachat dataset.
Training Details
The model underwent 20,000 training steps with a learning rate of 0.0001 and a total batch size of 256 across 8 GPUs. The training utilized an Adam optimizer and a linear learning rate scheduler with 500 warmup steps. The final validation loss achieved was 1.0042.
Frameworks Used
- Transformers 4.40.1
- Pytorch 2.2.2+cu121
- Datasets 2.19.0
- Tokenizers 0.19.1
Potential Use Cases
Given its fine-tuning on a chat-oriented dataset, this model is likely suitable for:
- General-purpose conversational agents
- Dialogue generation
- Instruction following in chat-based applications