lik07/Hercules-Stheno-v1

TEXT GENERATIONPricing:Input $0.37 / Cached $0.074 / Output $0.38Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Jun 15, 2024Architecture:Transformer Featherless Exclusive Cold

lik07/Hercules-Stheno-v1 is an 8 billion parameter language model created by lik07, resulting from a DARE TIES merge of Locutusque/Llama-3-Hercules-5.0-8B and Sao10K/L3-8B-Stheno-v3.2. This model leverages the strengths of its constituent models, offering a balanced performance profile for general language tasks. Its merge architecture aims to combine diverse capabilities from its Llama-3 based components.

Loading preview...

Model Overview

lik07/Hercules-Stheno-v1 is an 8 billion parameter language model developed by lik07. It is a merged model, combining the capabilities of two distinct pre-trained models using the DARE TIES (DARE with TIES-merging) method. This approach is designed to integrate the strengths of multiple models into a single, more robust entity.

Merge Details

This model was created by merging:

The DARE TIES merge method, as described in relevant research papers, allows for a strategic combination of model weights, aiming to preserve and enhance performance across various tasks. The configuration used for this merge specified a density of 0.5 and a weight of 1.0 for both constituent models, with int8_mask enabled and bfloat16 dtype.

Key Characteristics

  • Architecture: Based on the Llama-3 family, inheriting its foundational capabilities.
  • Parameter Count: 8 billion parameters, suitable for a range of applications requiring efficient inference.
  • Merge Method: Utilizes the DARE TIES method, a technique known for effectively combining models.

Potential Use Cases

Given its merged nature and Llama-3 base, Hercules-Stheno-v1 is likely suitable for:

  • General text generation and understanding tasks.
  • Applications benefiting from a blend of capabilities from its Hercules and Stheno components.
  • Scenarios where an 8B parameter model offers a good balance between performance and computational efficiency.