Domeoneyyddkk/qwen2.5-coder-merged

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 7, 2026Architecture:Transformer Featherless Exclusive Cold

Domeoneyyddkk/qwen2.5-coder-merged is a 7.6 billion parameter language model, merged from Qwen/Qwen2.5-Coder-7B-Instruct and Qwen/Qwen2.5-7B-Instruct using the DARE TIES method. With a 32768-token context length, this model is designed to leverage the strengths of both its base models, likely enhancing its capabilities in code generation and general instruction following. It aims to provide a balanced performance for tasks requiring both coding proficiency and broader language understanding.

Loading preview...

Model Overview

This model, Domeoneyyddkk/qwen2.5-coder-merged, is a 7.6 billion parameter language model created by merging existing pre-trained models using the mergekit tool.

Merge Details

The model was constructed using the DARE TIES merge method. The primary base model for this merge was Qwen/Qwen2.5-Coder-7B-Instruct, which was combined with Qwen/Qwen2.5-7B-Instruct.

Key Capabilities

  • Enhanced Code Generation: By incorporating Qwen2.5-Coder-7B-Instruct, the merged model is expected to retain and potentially improve its capabilities in understanding and generating code.
  • General Instruction Following: The inclusion of Qwen2.5-7B-Instruct contributes to robust general instruction-following abilities, making it versatile for various NLP tasks.
  • Balanced Performance: The DARE TIES method, with specific weighting (0.65 for Coder and 0.35 for Instruct), aims to create a model that balances specialized coding prowess with broad language understanding.

When to Use This Model

  • Code-centric Applications: Ideal for tasks requiring strong code generation, completion, and explanation.
  • Mixed-Task Environments: Suitable for use cases that involve both programming-related queries and general conversational or instructional prompts.
  • Research and Experimentation: Developers interested in exploring the effects of model merging on Qwen2.5 architectures for specific performance profiles.