trinhkhng/della_Merged_Qwen2-0.5B_0.5
The trinhkhng/della_Merged_Qwen2-0.5B_0.5 model is a 0.5 billion parameter language model based on the Qwen2 architecture, created by trinhkhng. It was developed using the DELLA merge method, combining a base Qwen2-0.5B model with a debiased version. This model is specifically designed to integrate debiasing characteristics into a compact language model while maintaining a 32768 token context length.
Loading preview...
Overview
This model, trinhkhng/della_Merged_Qwen2-0.5B_0.5, is a 0.5 billion parameter language model built upon the Qwen2 architecture. It was created by trinhkhng using the DELLA merge method, a technique designed to combine pre-trained language models. The merging process specifically integrated a debiased version of Qwen2-0.5B with a standard Qwen2-0.5B base model.
Key Capabilities
- DELLA Merge Method: Utilizes the DELLA (Density-based Layer-wise Linear Averaging) merge method, as detailed in the arXiv paper, to combine model characteristics.
- Debiasing Integration: Incorporates a debiased Qwen2-0.5B model, suggesting an aim to mitigate biases present in the base model.
- Compact Size: At 0.5 billion parameters, it offers a relatively small footprint for deployment.
- Extended Context Length: Supports a substantial context window of 32768 tokens, allowing for processing longer inputs.
Good For
- Experiments with model merging techniques, particularly DELLA.
- Applications requiring a compact language model with integrated debiasing efforts.
- Scenarios where a 32K context length is beneficial for a small model.