trinhkhng/della_Merged_Qwen2-0.5B_0.5

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 6, 2026Architecture:Transformer Featherless Exclusive Cold

The trinhkhng/della_Merged_Qwen2-0.5B_0.5 model is a 0.5 billion parameter language model based on the Qwen2 architecture, created by trinhkhng. It was developed using the DELLA merge method, combining a base Qwen2-0.5B model with a debiased version. This model is specifically designed to integrate debiasing characteristics into a compact language model while maintaining a 32768 token context length.

Loading preview...

Overview

This model, trinhkhng/della_Merged_Qwen2-0.5B_0.5, is a 0.5 billion parameter language model built upon the Qwen2 architecture. It was created by trinhkhng using the DELLA merge method, a technique designed to combine pre-trained language models. The merging process specifically integrated a debiased version of Qwen2-0.5B with a standard Qwen2-0.5B base model.

Key Capabilities

  • DELLA Merge Method: Utilizes the DELLA (Density-based Layer-wise Linear Averaging) merge method, as detailed in the arXiv paper, to combine model characteristics.
  • Debiasing Integration: Incorporates a debiased Qwen2-0.5B model, suggesting an aim to mitigate biases present in the base model.
  • Compact Size: At 0.5 billion parameters, it offers a relatively small footprint for deployment.
  • Extended Context Length: Supports a substantial context window of 32768 tokens, allowing for processing longer inputs.

Good For

  • Experiments with model merging techniques, particularly DELLA.
  • Applications requiring a compact language model with integrated debiasing efforts.
  • Scenarios where a 32K context length is beneficial for a small model.