MiniLLM/teacher-Llama-13B
MiniLLM/teacher-Llama-13B is a 13 billion parameter Llama model developed by MiniLLM, specifically fine-tuned on the databricks-dolly-15k dataset. This model is primarily designed to serve as a teacher model for knowledge distillation within the MiniLLM series, demonstrating its utility in guiding smaller language models. With a 4096-token context length, it specializes in providing high-quality supervised signals for training other LLMs.
Loading preview...
Overview
MiniLLM/teacher-Llama-13B is a 13 billion parameter language model built upon the Llama architecture. Developed by MiniLLM, this model has undergone supervised fine-tuning using the databricks-dolly-15k dataset. Its primary role is to function as a "teacher" model, providing high-quality supervision for the knowledge distillation process within the broader MiniLLM series. This makes it a foundational component for developing more efficient and smaller language models.
Key Capabilities
- Supervised Fine-Tuning: Leverages the
databricks-dolly-15kdataset for instruction-following capabilities. - Knowledge Distillation: Designed specifically to act as a teacher model, guiding the training of smaller student models.
- Llama Architecture: Benefits from the robust and widely recognized Llama model family.
Good For
- Research in Knowledge Distillation: Ideal for researchers and developers exploring methods to transfer knowledge from large models to smaller, more efficient ones.
- Generating Training Data: Can be used to create high-quality, instruction-tuned responses for training other language models.
- Foundation for MiniLLM Series: Serves as a core component for projects within the MiniLLM ecosystem requiring a strong teacher signal.