Sakalti/mont-normal

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Oct 24, 2024Architecture:Transformer Featherless Exclusive Cold

Sakalti/mont-normal is a 0.5 billion parameter language model created by Sakalti, merged using the TIES method with Qwen/Qwen2.5-0.5B-Instruct as its base. This model leverages specific layer ranges from the Qwen2.5-0.5B-Instruct model, applying varying parameter values for self-attention and MLP filters. It is designed for general language tasks, inheriting the foundational capabilities of its Qwen2.5 base within a compact 0.5B parameter size and supporting a 32768 token context length.

Loading preview...

Model Overview

Sakalti/mont-normal is a compact 0.5 billion parameter language model developed by Sakalti. It was created through a merging process using mergekit, specifically employing the TIES merge method. The foundational model for this merge is Qwen/Qwen2.5-0.5B-Instruct, indicating its base capabilities are derived from the Qwen2.5 family.

Merge Details

The model's unique configuration involves merging different layer ranges of the Qwen/Qwen2.5-0.5B-Instruct model itself. This process applied specific parameter adjustments, particularly for self-attention and MLP filters, with varying values (0.1, 0.3, 0.5, 0.7, 0.9) to fine-tune the merged model's characteristics. The merge was performed using bfloat16 precision.

Key Characteristics

  • Base Model: Built upon Qwen/Qwen2.5-0.5B-Instruct.
  • Merge Method: Utilizes the TIES (Trimming and Expanding the Search Space) method for combining model components.
  • Parameter Count: A compact 0.5 billion parameters, making it suitable for resource-constrained environments.
  • Context Length: Supports a substantial context window of 32768 tokens.

Use Cases

This model is suitable for applications requiring a small, efficient language model with capabilities inherited from the Qwen2.5-Instruct series. Its compact size and efficient merging technique suggest potential for tasks where computational resources are limited but a capable language understanding and generation model is needed.