expertai/SLIMER
expertai/SLIMER is a 7 billion parameter language model developed by expertai, specifically instruction-tuned for zero-shot Named Entity Recognition (NER) in English. Built on a LLaMA-2-chat backbone, it excels at identifying never-before-seen entity tags by leveraging prompts enriched with explicit definitions and guidelines for the target entities. This approach allows SLIMER to perform comparably to state-of-the-art models on out-of-distribution NER tasks, particularly demonstrating strong performance on specialized financial entities.
Loading preview...
SLIMER: Zero-Shot NER with Enriched Prompts
expertai/SLIMER (Show Less Instruct More Entity Recognition) is a 7 billion parameter model based on the LLaMA-2-chat architecture, specifically designed for zero-shot Named Entity Recognition (NER) in English. Unlike traditional approaches that rely on extensive fine-tuning across thousands of entity classes, SLIMER is instructed on a reduced number of samples.
Key Capabilities and Differentiators
- Zero-Shot NER with Definitions and Guidelines: SLIMER's core innovation lies in its prompting strategy. It tackles novel entity tags by incorporating a DEFINITION and GUIDELINES for the target Named Entity directly into the prompt. This allows it to understand and extract entities it has not explicitly seen during training.
- Performance on Unseen Entities: The model demonstrates strong performance on out-of-distribution (OOD) input domains, including specialized financial entities (BUSTER dataset), where it often outperforms other state-of-the-art models. This is attributed to its lighter instruction tuning methodology and the effective use of definitional prompts.
- Efficient Training: SLIMER achieves competitive results while being trained on a significantly reduced number of samples and a less overlapping set of NE tags compared to test sets, making its instruction tuning more efficient.
Use Cases
SLIMER is particularly well-suited for:
- Extracting novel or domain-specific entities where pre-trained models might struggle due to lack of specific training data.
- Applications requiring flexible NER that can adapt to new entity types on the fly without extensive re-training.
- Tasks where precise entity definitions and extraction rules can be provided to guide the model, such as in legal, medical, or financial text analysis.