ai4bharat/hercule-bn

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 18, 2024License:mitArchitecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Hercule-bn is an 8 billion parameter cross-lingual evaluation model developed by AI4Bharat, fine-tuned on Llama-3.1-8B-Instruct with a 32768 token context length. It is designed to assess multilingual Large Language Models (LLMs) by using English reference responses to score multilingual outputs, particularly excelling in low-resource scenarios and supporting zero-shot evaluations on unseen languages. The model provides feedback and scores on a 1-5 scale, demonstrating better alignment with human judgments compared to zero-shot evaluations by proprietary models like GPT-4 on the RECON test set.

Loading preview...

Hercule-bn: Cross-Lingual LLM Evaluation

Hercule-bn is an 8 billion parameter evaluation model developed by AI4Bharat, specifically fine-tuned on Llama-3.1-8B-Instruct. It is part of the broader CIA Suite, designed to address the challenges of evaluating multilingual Large Language Models (LLMs) by leveraging English reference responses to score outputs in various languages.

Key Capabilities & Features

  • Cross-Lingual Evaluation: Scores multilingual LLM outputs using English reference answers and evaluation criteria.
  • Human Alignment: Demonstrates better alignment with human judgments on the RECON test set compared to zero-shot evaluations by models like GPT-4.
  • Low-Resource & Zero-Shot: Excels in low-resource language scenarios and supports zero-shot evaluations for unseen languages.
  • Reference-Based Scoring: Provides detailed feedback and a 1-5 score based on a given rubric, instruction, response, and reference answer.
  • Efficient Fine-tuning: Utilizes lightweight fine-tuning methods (like LoRA) for efficient multilingual evaluation.

Use Cases & Differentiators

Hercule-bn is ideal for developers and researchers needing to objectively evaluate the performance of multilingual LLMs, especially for Bengali. Its primary differentiator is its ability to provide robust, reference-guided evaluation across languages, offering a more reliable assessment than generic zero-shot methods. The model's fine-tuning on the INTEL dataset and its use of a Prometheus 2-like prompt format contribute to its effectiveness in generating actionable feedback and scores. It is particularly useful for assessing LLMs in contexts where human evaluation is costly or impractical, providing a scalable solution for quality control and performance benchmarking.