trinhkhng/della_Merged_Qwen2-0.5B_0.3

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 6, 2026Architecture:Transformer Featherless Exclusive Cold

trinhkhng/della_Merged_Qwen2-0.5B_0.3 is a 0.5 billion parameter language model created by trinhkhng, merged from a Qwen2-0.5B base using the DELLA merge method. This model specifically incorporates a debiased version of Qwen2-0.5B, suggesting an optimization for reduced bias. With a 32768 token context length, it is suitable for applications requiring smaller, efficient models with potentially improved fairness characteristics.

Loading preview...

Model Overview

trinhkhng/della_Merged_Qwen2-0.5B_0.3 is a 0.5 billion parameter language model derived from the Qwen2-0.5B architecture. This model was created by trinhkhng using the DELLA merge method, a technique designed for combining pre-trained language models. A key characteristic of this merge is the inclusion of a debiased Qwen2-0.5B model, indicating an effort to mitigate biases present in the base model.

Key Characteristics

  • Architecture: Based on the Qwen2-0.5B model.
  • Parameter Count: 0.5 billion parameters, making it a relatively small and efficient model.
  • Context Length: Supports a substantial context window of 32768 tokens.
  • Merge Method: Utilizes the DELLA merge method, which allows for combining models with specific configurations.
  • Bias Mitigation: Incorporates a debiased version of the Qwen2-0.5B model, suggesting potential improvements in fairness.

Potential Use Cases

This model is well-suited for applications where:

  • Computational resources are limited, benefiting from its smaller parameter count.
  • Long context understanding is required, thanks to its 32768 token context length.
  • Reduced bias in language generation is a priority, due to the inclusion of a debiased component.