mergekit-community/Berry-Spark-7B

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 12, 2024Architecture:Transformer Featherless Exclusive Cold

Berry-Spark-7B is a 7.6 billion parameter language model created by mergekit-community, formed by merging arcee-ai/Arcee-Spark and ChaoticNeutrals/Very_Berry_Qwen2_7B using the SLERP method. This model leverages a 32768 token context length, combining the strengths of its base models to offer enhanced general-purpose language capabilities. It is designed for applications requiring a robust and versatile 7B class model.

Loading preview...

Overview

Berry-Spark-7B is a 7.6 billion parameter language model developed by mergekit-community. It was created using the SLERP merge method to combine two distinct base models: arcee-ai/Arcee-Spark and ChaoticNeutrals/Very_Berry_Qwen2_7B. This merging process aims to integrate the strengths of both components into a single, more capable model.

Merge Details

The merge utilized a specific configuration, applying varying weights across different layers and attention mechanisms. The self_attn layers were weighted with a complex t parameter array, as were the mlp layers, indicating a fine-tuned blend of the source models' characteristics. The base model for the merge was arcee-ai/Arcee-Spark, with the final output dtype set to bfloat16.

Key Characteristics

  • Parameter Count: 7.6 billion parameters, offering a balance between performance and computational efficiency.
  • Context Length: Supports a substantial context window of 32768 tokens, enabling processing of longer inputs and maintaining coherence over extended conversations or documents.
  • Merge Method: Employs the SLERP (Spherical Linear Interpolation) method, known for producing stable and effective merges of neural network weights.

Potential Use Cases

Given its merged architecture and substantial context, Berry-Spark-7B is suitable for a variety of general-purpose language tasks, including:

  • Text generation and completion
  • Summarization of long documents
  • Conversational AI and chatbots
  • Code generation and understanding (depending on the base models' capabilities)