numind/NuExtract-1.5-tiny

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 26, 2024License:mitArchitecture:Transformer0.0K Open Weights Featherless Exclusive Warm

NuExtract-1.5-tiny by NuMind is a 0.5 billion parameter language model, fine-tuned from Qwen2.5-0.5B, specifically designed for structured information extraction from long documents. It excels at extracting data into a JSON template across multiple languages including English, French, Spanish, German, Portuguese, and Italian. The model prioritizes pure extraction, ensuring generated text is present in the original source, and supports a 32768 token context length.

Loading preview...

NuExtract-1.5-tiny: Specialized Information Extraction

NuExtract-1.5-tiny, developed by NuMind, is a 0.5 billion parameter model fine-tuned from Qwen2.5-0.5B. Its core purpose is structured information extraction from text, supporting a context length of 32768 tokens. The model is trained on a private, high-quality dataset to accurately extract data into a user-defined JSON template.

Key Capabilities

  • Multilingual Extraction: Supports information extraction in English, French, Spanish, German, Portuguese, and Italian.
  • Pure Extraction Focus: Designed to ensure that all generated text is directly present in the original input text, minimizing hallucination for extraction tasks.
  • Long Document Support: Capable of processing extensive documents, making it suitable for complex data extraction scenarios.
  • Template-Driven Output: Users provide a JSON template to guide the extraction process, ensuring structured and consistent output.
  • Optimized for Low Temperature: Recommended for use with a temperature setting at or near 0 for optimal extraction accuracy.

Good For

  • Automated Data Extraction: Ideal for tasks requiring the structured extraction of specific entities or information from unstructured text.
  • Multilingual Content Processing: Suitable for businesses operating with documents in the supported European languages.
  • Reducing Hallucination: Its design prioritizes direct extraction, making it reliable for applications where factual accuracy from the source text is paramount.

NuMind also offers a larger 3.8B parameter version, NuExtract-v1.5, based on Phi-3.5-mini-instruct, for more demanding extraction needs.

Popular Sampler Settings

Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.

temperature
top_p
top_k
frequency_penalty
presence_penalty
repetition_penalty
min_p