MiniLLM/SFT-Llama-7B

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7BQuant:FP8Context Size:4kPublished:Sep 26, 2024License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

MiniLLM/SFT-Llama-7B is a 7 billion parameter Llama-based model developed by MiniLLM, specifically fine-tuned using supervised learning on the databricks-dolly-15k dataset. This model serves as a foundational baseline for the MiniLLM project, demonstrating capabilities derived from instruction-following data. It is designed for general language understanding and generation tasks, particularly those benefiting from instruction-tuned responses.

Loading preview...

Overview

MiniLLM/SFT-Llama-7B is a 7 billion parameter language model built upon the Llama architecture. Developed by MiniLLM, this model has undergone supervised fine-tuning (SFT) using the databricks-dolly-15k dataset. It functions as a key baseline model within the broader MiniLLM project, which explores knowledge distillation techniques for large language models.

Key Capabilities

  • Instruction Following: Fine-tuned on a dataset designed for instruction-following, enabling it to respond to a variety of prompts and instructions.
  • General Language Tasks: Suitable for a range of natural language processing tasks, including text generation, summarization, and question answering, based on its supervised fine-tuning.
  • Baseline Model: Serves as a reference point for evaluating other models within the MiniLLM series, such as KD-Llama-7B and SeqKD-Llama-7B, which incorporate knowledge distillation.

Good For

  • Developers seeking a Llama-7B based model with instruction-following capabilities.
  • Researchers interested in baselines for knowledge distillation experiments.
  • Applications requiring a moderately sized model for general text generation and understanding.