Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.7

TEXT GENERATIONConcurrent Unit Cost:1Model Size:14.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Mar 1, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.7 is a 14.8 billion parameter language model based on the Qwen2.5 architecture, created by Lunzima using the Model Stock merge method. This model integrates multiple specialized 14B models, including Lamarck-14B-v0.7-Fusion and Qwenvergence-14B-v11, to enhance its overall capabilities. With a context length of 32768 tokens, it is designed for broad application across various natural language processing tasks, leveraging the strengths of its diverse merged components.

Loading preview...

Model Overview

Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.7 is a 14.8 billion parameter language model built upon the Qwen2.5 architecture. Developed by Lunzima, this model was created using the advanced Model Stock merge method, which combines the strengths of several pre-trained language models into a single, more robust entity. The base model for this fusion was Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.

Key Merged Components

This model is a sophisticated merge of seven distinct 14B parameter models, each contributing unique characteristics to the final fusion. Notable components include:

  • sometimesanotion/Lamarck-14B-v0.7-Fusion
  • sometimesanotion/Qwenvergence-14B-v11
  • prithivMLmods/Messier-Opus-14B-Elite7
  • jpacifico/Chocolatine-2-14B-Instruct-v2.0b3
  • prithivMLmods/Equuleus-Opus-14B-Exp
  • Sakalti/Saka-14B

Technical Specifications

The merging process utilized mergekit and was configured to use bfloat16 data type for efficiency. The tokenizer source was set to union, ensuring comprehensive tokenization capabilities derived from the merged models. The model supports a substantial context length of 32768 tokens, making it suitable for processing extensive inputs and generating detailed outputs. The int8_mask parameter was enabled during the merge, indicating potential optimizations for quantization.