opencompass/CompassJudger-1-1.5B-Instruct

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Oct 16, 2024License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The opencompass/CompassJudger-1-1.5B-Instruct is a 1.5 billion parameter instruction-tuned model developed by Opencompass, designed as an all-in-one judge model. It excels in comprehensive evaluation methods, including scoring and comparison, and can output detailed assessment reviews in specified formats. This model also functions as a versatile general instruction model, supporting inference acceleration methods like vLLM and LMdeploy, making it suitable for diverse evaluation datasets and daily tasks.

Loading preview...

Overview

Opencompass's CompassJudger-1 series, including this 1.5 billion parameter instruction-tuned model, is an all-in-one judge model designed for comprehensive evaluation tasks. It distinguishes itself by offering multiple evaluation methods, such as scoring and comparison, and can generate detailed assessment reviews in a specified format. Beyond its primary judging capabilities, CompassJudger-1 also functions as a versatile general instruction model, capable of handling typical daily tasks.

Key Capabilities

  • Comprehensive Evaluation: Performs various evaluation methods, including scoring, comparison, and detailed assessment feedback.
  • Formatted Output: Supports generating evaluation results in specific, user-defined formats for easier analysis.
  • Versatility: Acts as a universal instruction model for general tasks, in addition to its specialized judging functions.
  • Inference Acceleration: Compatible with model inference acceleration methods like vLLM and LMdeploy.

Use Cases

  • Reward Model: Can be used as a reward model in scenarios requiring comparative judgments between different AI assistant responses.
  • Point-wise Judging: Evaluates AI assistant responses against specific criteria, providing detailed critiques and scores across multiple dimensions.
  • Response Critique: Offers constructive feedback and suggestions for improving AI-generated content based on user requirements.
  • Subjective Dataset Evaluation: Integrates with OpenCompass for evaluating subjective datasets, facilitating standardized assessment of various models.

Opencompass has also established JudgerBench, a benchmark to standardize the evaluation of judging models, with a dedicated leaderboard available.