ai4bharat/hercule-de

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 18, 2024License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

Hercule-DE is an 8 billion parameter cross-lingual evaluation model developed by AI4Bharat, fine-tuned on Llama-3.1-8B-Instruct. It specializes in assessing multilingual Large Language Models by comparing their outputs to English reference responses, providing a 1-5 score and feedback. This model demonstrates strong alignment with human judgments, particularly in low-resource scenarios, and supports zero-shot evaluation for unseen languages. Its primary use is for automated, reference-based quality assessment of LLM responses in German.

Loading preview...

Hercule-DE: Cross-Lingual LLM Evaluation Model

Hercule-DE is an 8 billion parameter evaluation model developed by AI4Bharat, specifically designed to assess the quality of multilingual Large Language Models (LLMs). It is fine-tuned on the Llama-3.1-8B-Instruct base model using the INTEL dataset, focusing on German language evaluation.

Key Capabilities and Features

  • Cross-Lingual Evaluation: Hercule-DE evaluates multilingual LLM responses by comparing them against English reference answers, providing a standardized scoring mechanism.
  • Human Alignment: It demonstrates better alignment with human judgments compared to zero-shot evaluations by proprietary models like GPT-4, particularly on the RECON test set.
  • Low-Resource Performance: The model excels in evaluating LLMs in low-resource language scenarios.
  • Zero-Shot Evaluation: Supports zero-shot evaluation for languages it has not been explicitly trained on.
  • Reference-Based Scoring: Provides detailed feedback and a score on a 1-5 scale, based on a provided rubric, instruction, response, and reference answer.
  • Efficient Fine-tuning: Highlights the effectiveness of lightweight fine-tuning methods like LoRA for multilingual evaluation.

Use Cases and Differentiators

Hercule-DE is ideal for developers and researchers needing an automated, robust method to evaluate the quality of LLM outputs in German. Its ability to use English references for scoring multilingual responses simplifies the evaluation pipeline. The model's strong performance in low-resource settings and its alignment with human judgment make it a valuable tool for quality assurance and development of multilingual LLMs. It utilizes a prompt format similar to Prometheus 2 for direct assessment.