opencompass/CompassJudger-1-1.5B-Instruct
The opencompass/CompassJudger-1-1.5B-Instruct is a 1.5 billion parameter instruction-tuned model developed by Opencompass, designed as an all-in-one judge model. It excels in comprehensive evaluation methods, including scoring and comparison, and can output detailed assessment reviews in specified formats. This model also functions as a versatile general instruction model, supporting inference acceleration methods like vLLM and LMdeploy, making it suitable for diverse evaluation datasets and daily tasks.
Loading preview...
Overview
Opencompass's CompassJudger-1 series, including this 1.5 billion parameter instruction-tuned model, is an all-in-one judge model designed for comprehensive evaluation tasks. It distinguishes itself by offering multiple evaluation methods, such as scoring and comparison, and can generate detailed assessment reviews in a specified format. Beyond its primary judging capabilities, CompassJudger-1 also functions as a versatile general instruction model, capable of handling typical daily tasks.
Key Capabilities
- Comprehensive Evaluation: Performs various evaluation methods, including scoring, comparison, and detailed assessment feedback.
- Formatted Output: Supports generating evaluation results in specific, user-defined formats for easier analysis.
- Versatility: Acts as a universal instruction model for general tasks, in addition to its specialized judging functions.
- Inference Acceleration: Compatible with model inference acceleration methods like vLLM and LMdeploy.
Use Cases
- Reward Model: Can be used as a reward model in scenarios requiring comparative judgments between different AI assistant responses.
- Point-wise Judging: Evaluates AI assistant responses against specific criteria, providing detailed critiques and scores across multiple dimensions.
- Response Critique: Offers constructive feedback and suggestions for improving AI-generated content based on user requirements.
- Subjective Dataset Evaluation: Integrates with OpenCompass for evaluating subjective datasets, facilitating standardized assessment of various models.
Opencompass has also established JudgerBench, a benchmark to standardize the evaluation of judging models, with a dedicated leaderboard available.