MergekitCloud/mergekit-103

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 6, 2026Architecture:Transformer Featherless Exclusive Cold

MergekitCloud/mergekit-103 is an 8-billion parameter language model created by MergekitCloud using the Linear DARE merge method, based on Qwen/Qwen3-8B. It integrates ertghiu256/qwen3-8b-code-reasoning and deepseek-ai/DeepSeek-R1-0528-Qwen3-8B, suggesting an optimization for code reasoning and general performance. This model is designed for tasks requiring a blend of coding capabilities and robust language understanding within a 32768-token context window.

Loading preview...

Overview

MergekitCloud/mergekit-103 is an 8-billion parameter language model developed by MergekitCloud. It was created using the Linear DARE merge method, building upon the Qwen/Qwen3-8B base model. This approach combines the strengths of multiple specialized models to enhance overall performance.

Key Capabilities

  • Enhanced Code Reasoning: Integrates ertghiu256/qwen3-8b-code-reasoning, suggesting improved capabilities in understanding and generating code-related logic.
  • Robust Language Understanding: Incorporates deepseek-ai/DeepSeek-R1-0528-Qwen3-8B, contributing to general language proficiency and reasoning.
  • Efficient Merging: Utilizes the Linear DARE method, which is designed to effectively combine pre-trained models while maintaining performance.
  • Large Context Window: Supports a context length of 32768 tokens, suitable for processing extensive inputs.

Good for

  • Applications requiring a balance of general language understanding and specialized code reasoning.
  • Tasks that benefit from a large context window, such as summarizing long documents or complex codebases.
  • Developers looking for a merged model that leverages the strengths of Qwen3-8B and fine-tuned components for specific domains.