lik07/Hermes-Sthero-v1

TEXT GENERATIONPricing:Input $0.37 / Cached $0.074 / Output $0.38Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Jun 16, 2024Architecture:Transformer Featherless Exclusive Cold

lik07/Hermes-Sthero-v1 is an 8 billion parameter language model created by lik07, merged using the DARE TIES method from NousResearch/Hermes-2-Pro-Llama-3-8B and Sao10K/L3-8B-Stheno-v3.2. This model combines the strengths of its base components, offering a balanced performance for general language tasks with an 8192-token context length. It is suitable for applications requiring a robust 8B parameter model derived from established Llama-3 architectures.

Loading preview...

Model Overview

lik07/Hermes-Sthero-v1 is an 8 billion parameter language model, a product of merging two pre-trained models: NousResearch/Hermes-2-Pro-Llama-3-8B and Sao10K/L3-8B-Stheno-v3.2. This merge was performed using the DARE (Dropout-based Adaptive Reweighting of Embeddings) TIES (Task-Independent Embedding Space) method, a technique designed to combine the capabilities of multiple models effectively.

Key Characteristics

  • Architecture: Based on the Llama-3 family, leveraging the strengths of its constituent models.
  • Merge Method: Utilizes the DARE TIES method, which involves reweighting and dropping out parameters during the merge process to optimize performance.
  • Base Models: Integrates features from both the NousResearch Hermes-2-Pro-Llama-3-8B, known for its instruction-following capabilities, and Sao10K/L3-8B-Stheno-v3.2.
  • Context Length: Supports an 8192-token context window, enabling processing of moderately long inputs.

Intended Use Cases

This model is suitable for a variety of general-purpose language tasks where an 8 billion parameter model offers a good balance of performance and computational efficiency. Its merged nature suggests a broad applicability, potentially excelling in areas where its base models demonstrated proficiency, such as instruction following and general text generation.