MiniLLM/teacher-Llama-13B

TEXT GENERATIONPricing:Input $1.5 / Output $2.1Concurrent Unit Cost:1Model Size:13BQuant:FP8Context Size:4kPublished:Sep 26, 2024License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

MiniLLM/teacher-Llama-13B is a 13 billion parameter Llama model developed by MiniLLM, specifically fine-tuned on the databricks-dolly-15k dataset. This model is primarily designed to serve as a teacher model for knowledge distillation within the MiniLLM series, demonstrating its utility in guiding smaller language models. With a 4096-token context length, it specializes in providing high-quality supervised signals for training other LLMs.

Loading preview...

Overview

MiniLLM/teacher-Llama-13B is a 13 billion parameter language model built upon the Llama architecture. Developed by MiniLLM, this model has undergone supervised fine-tuning using the databricks-dolly-15k dataset. Its primary role is to function as a "teacher" model, providing high-quality supervision for the knowledge distillation process within the broader MiniLLM series. This makes it a foundational component for developing more efficient and smaller language models.

Key Capabilities

  • Supervised Fine-Tuning: Leverages the databricks-dolly-15k dataset for instruction-following capabilities.
  • Knowledge Distillation: Designed specifically to act as a teacher model, guiding the training of smaller student models.
  • Llama Architecture: Benefits from the robust and widely recognized Llama model family.

Good For

  • Research in Knowledge Distillation: Ideal for researchers and developers exploring methods to transfer knowledge from large models to smaller, more efficient ones.
  • Generating Training Data: Can be used to create high-quality, instruction-tuned responses for training other language models.
  • Foundation for MiniLLM Series: Serves as a core component for projects within the MiniLLM ecosystem requiring a strong teacher signal.