MergekitCloud/mergekit-102

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 6, 2026Architecture:Transformer Featherless Exclusive Cold

MergekitCloud/mergekit-102 is an 8 billion parameter language model created by MergekitCloud using the linear merge method. This model combines Qwen3-8B, ertghiu256/qwen3-8b-code-reasoning, and deepseek-ai/DeepSeek-R1-0528-Qwen3-8B, leveraging their strengths. With a 32768 token context length, it is designed for enhanced performance across general language tasks and specialized code reasoning applications.

Loading preview...

Overview

MergekitCloud/mergekit-102 is an 8 billion parameter language model developed by MergekitCloud. It was constructed using the linear merge method from the mergekit framework, combining the strengths of several pre-trained models. This approach aims to create a robust model by integrating diverse capabilities from its constituent parts.

Key Capabilities

  • Enhanced General Language Understanding: Incorporates the base capabilities of Qwen3-8B for broad language tasks.
  • Specialized Code Reasoning: Benefits from the inclusion of ertghiu256/qwen3-8b-code-reasoning, suggesting improved performance in code-related tasks and logical reasoning.
  • Integrated Strengths: Leverages contributions from deepseek-ai/DeepSeek-R1-0528-Qwen3-8B to further refine its overall performance.

Merge Details

The model was created by merging three distinct models with specific weighting parameters:

  • Qwen/Qwen3-8B (weight: 0.34)
  • ertghiu256/qwen3-8b-code-reasoning (weight: 0.55)
  • deepseek-ai/DeepSeek-R1-0528-Qwen3-8B (weight: 0.33)
    The merge process normalized these weights and used bfloat16 for the data type, with Qwen/Qwen3-8B serving as the tokenizer source. The model supports an auto chat template.

Good For

This model is suitable for applications requiring a balance of general language understanding and specific strengths in areas like code reasoning, making it a versatile choice for developers looking for a merged 8B parameter model.