appvoid/palmer-004

Hugging Face
TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.1BQuant:BF16Context Size:2kPublished:Jun 2, 2024License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Warm

The appvoid/palmer-004 is a 1.1 billion parameter language model developed by appvoid, featuring a 2048-token context length. This model demonstrates improved overall performance across various benchmarks, including MMLU, ARC-C, HellaSwag, and PIQA, making it suitable for general language understanding tasks. It is designed to respond to queries without specific prompting and can be further fine-tuned for specialized applications.

Loading preview...

appvoid/palmer-004: An Overview

appvoid/palmer-004 is a 1.1 billion parameter language model with a 2048-token context window, developed by appvoid. This model represents an iteration with enhanced overall performance compared to its predecessor, palmer-004-old, though it exhibits a slight degradation on the Winogrande benchmark.

Key Capabilities & Performance

The model shows improved scores across several benchmarks, indicating its general language understanding and reasoning abilities:

  • MMLU: 0.2661
  • ARC-C: 0.3490
  • HellaSwag: 0.6173
  • PIQA: 0.7481
  • Winogrande: 0.6417 (slightly lower than palmer-004-old's 0.6511)
  • Average Score: 0.5244

Notably, palmer-004 surpasses tinyllama-3t and palmer-004-old in average performance. The model is inherently biased to respond without requiring specific prompts, offering flexibility for various applications.

When to Use This Model

  • General Language Tasks: Its improved benchmark scores suggest suitability for a range of common NLP tasks.
  • Fine-tuning: The model is designed to be easily fine-tuned for specific use cases, allowing developers to adapt its behavior to their needs.
  • Short Context Applications: With a 2048-token context size, it is well-suited for tasks that do not require extensive context understanding, such as short-form content generation or quick question-answering.