ai4bharat/hercule-ur
Hercule-ur is an 8 billion parameter cross-lingual evaluation model developed by AI4Bharat, fine-tuned on Llama-3.1-8B-Instruct. It specializes in assessing multilingual Large Language Models by using English reference responses to score multilingual outputs, particularly for Urdu. This model demonstrates improved alignment with human judgments in low-resource scenarios and supports zero-shot evaluations on unseen languages, providing feedback and scores on a 1-5 scale.
Loading preview...
Hercule-ur: A Cross-Lingual Evaluation Model
Hercule-ur is an 8 billion parameter evaluation model developed by AI4Bharat, specifically designed for assessing multilingual Large Language Models (LLMs). It is fine-tuned on the Llama-3.1-8B-Instruct architecture using the INTEL dataset, with a focus on Urdu language evaluation. This model addresses the challenge of evaluating multilingual LLMs by leveraging English reference responses to score outputs in other languages.
Key Capabilities and Features
- Cross-Lingual Evaluation: Utilizes English reference answers and scoring rubrics to evaluate multilingual responses, providing a standardized assessment framework.
- Human Alignment: Demonstrates better alignment with human judgments compared to zero-shot evaluations by proprietary models like GPT-4 on the RECON test set.
- Low-Resource Performance: Excels particularly in scenarios involving low-resource languages, making it valuable for diverse linguistic contexts.
- Zero-Shot Evaluation: Supports zero-shot evaluation on unseen languages, enhancing its adaptability.
- Reference-Based Scoring: Provides detailed feedback and a score on a 1-5 scale, based on a given instruction, response, reference answer, and evaluation criteria.
- Efficient Fine-tuning: Highlights the effectiveness of lightweight fine-tuning methods like LoRA for efficient multilingual evaluation.
Use Cases and Differentiators
Hercule-ur is ideal for developers and researchers needing to objectively evaluate the quality of multilingual LLM outputs, especially for Urdu. Its ability to provide consistent, reference-guided scores makes it a robust tool for benchmarking and improving multilingual models. The model's prompt format requires an evaluation instruction, a response to evaluate, a reference answer, and a detailed scoring rubric, similar to the Prometheus 2 evaluation prompt. Further details and related models are available in the CIA Suite and its research paper.