prometheus-eval/prometheus-7b-v1.0

Hugging Face
TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7BQuant:FP8Context Size:4kPublished:Oct 12, 2023License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Warm

Prometheus-7b-v1.0 by KAIST AI is a 7 billion parameter language model based on Llama-2-Chat, fine-tuned on 100K feedback examples from the Feedback Collection dataset. It specializes in fine-grained evaluation of long-form LLM responses, outperforming GPT-3.5-Turbo and Llama-2-Chat 70B, and performing on par with GPT-4 on various benchmarks. This model is designed as a cost-effective alternative to GPT-4 for evaluating LLMs with customized criteria and can also function as a reward model for Reinforcement Learning from Human Feedback (RLHF).

Loading preview...

Overview

Prometheus-7b-v1.0 is a 7 billion parameter language model developed by KAIST AI, built upon the Llama-2-Chat architecture. It has been extensively fine-tuned using 100,000 feedback examples from the dedicated Feedback Collection dataset. This specialization enables Prometheus to excel in the nuanced evaluation of long-form responses generated by other large language models.

Key Capabilities

  • Fine-grained LLM Evaluation: Prometheus is specifically designed to provide detailed, criterion-based evaluations of LLM outputs, leveraging reference answers and customized score rubrics.
  • Cost-Effective Alternative to GPT-4: It offers a powerful and more economical solution for evaluation tasks that typically require models like GPT-4, matching its performance on various benchmarks.
  • Customizable Criteria: Users can define their own evaluation criteria (e.g., child readability, cultural sensitivity, creativity) through detailed score rubrics.
  • RLHF Reward Model: The model can be effectively utilized as a reward model within Reinforcement Learning from Human Feedback (RLHF) pipelines.
  • Performance: Outperforms GPT-3.5-Turbo and Llama-2-Chat 70B in evaluation tasks, achieving performance comparable to GPT-4.

When to Use This Model

  • Evaluating LLM Responses: Ideal for developers and researchers needing objective, detailed feedback on the quality of LLM-generated text, especially long-form content.
  • Custom Evaluation Metrics: When standard evaluation metrics are insufficient, and specific, custom criteria are required.
  • RLHF Applications: Suitable for integration into RLHF systems as a reward model to guide model training based on human-like feedback.
  • Resource-Constrained Evaluation: A strong choice for high-quality evaluation when the cost or accessibility of larger models like GPT-4 is a concern.