trinhkhng/della_Merged_Qwen2-0.5B_0.4

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 6, 2026Architecture:Transformer Featherless Exclusive Cold

trinhkhng/della_Merged_Qwen2-0.5B_0.4 is a 0.5 billion parameter language model based on the Qwen2 architecture, created by trinhkhng using the DELLA merge method. This model integrates a debiased Qwen2-0.5B variant into a Qwen2-0.5B base, focusing on specific parameter adjustments for its merged configuration. It is designed for applications requiring a compact yet specialized language model, leveraging its unique merging approach.

Loading preview...

Model Overview

trinhkhng/della_Merged_Qwen2-0.5B_0.4 is a 0.5 billion parameter language model, built upon the Qwen2-0.5B base architecture. This model was developed by trinhkhng using the DELLA merge method, a technique for combining pre-trained language models.

Merge Details

The model's unique characteristic lies in its merging process. It incorporates a debiased version of Qwen2-0.5B into the standard Qwen2-0.5B base. The merge was configured with specific parameters, including a density of 0.5, an epsilon of 0.1, and a weight of 1.0 for the debiased model. The overall merge also utilized int8_mask, normalize, and rescale parameters, with a lambda value of 0.4, indicating a tailored approach to model integration.

Key Characteristics

  • Architecture: Qwen2-0.5B base.
  • Parameter Count: 0.5 billion parameters.
  • Merge Method: Utilizes the DELLA method for combining models.
  • Component Models: Merges a debiased Qwen2-0.5B with a standard Qwen2-0.5B.
  • Configuration: Specific merge parameters applied for fine-grained control over the integration process.

Potential Use Cases

This model is suitable for applications where a compact, specialized language model is required, particularly in scenarios that might benefit from a debiased component integrated via the DELLA method. Its small size makes it efficient for deployment in resource-constrained environments or for tasks that do not require the scale of larger models.