kechengcode/Qwen2-5B-16Layers

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 22, 2024Architecture:Transformer Featherless Exclusive Cold

kechengcode/Qwen2-5B-16Layers is a 7.6 billion parameter language model derived from Qwen/Qwen2-7B, created by kechengcode through a passthrough merge method. This model specifically integrates selected layers from the base Qwen2-7B architecture, offering a reconfigured version of the original Qwen2 model. It is suitable for applications requiring a Qwen2-based model with a modified layer structure.

Loading preview...

Overview

kechengcode/Qwen2-5B-16Layers is a 7.6 billion parameter language model, which is a merged derivative of the Qwen/Qwen2-7B base model. This model was created using the mergekit tool with a passthrough merge method, specifically selecting and combining certain layer ranges from the original Qwen2-7B architecture.

Merge Details

This model is not a standalone training but a reconfiguration of an existing powerful model. The merge process involved:

  • Base Model: Qwen/Qwen2-7B
  • Merge Method: Passthrough, which directly incorporates specified layers.
  • Layer Selection: The merge specifically included layers [0, 10] and [22, 28] from the Qwen/Qwen2-7B model, resulting in a 16-layer structure.

Key Characteristics

  • Architecture: Based on the Qwen2 family, known for its strong performance across various language tasks.
  • Parameter Count: 7.6 billion parameters, making it a substantial model for diverse applications.
  • Context Length: Inherits a 32768 token context length, allowing for processing of long inputs.

Use Cases

This model is suitable for developers looking to leverage the Qwen2 architecture with a specific layer configuration. It can be used for general language understanding, generation, and other NLP tasks where the Qwen2-7B's capabilities are desired within this merged structure.