trinhkhng/nuslerp_Merged_Qwen2-0.5B_0.3

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 6, 2026Architecture:Transformer Featherless Exclusive Cold

The trinhkhng/nuslerp_Merged_Qwen2-0.5B_0.3 is a 0.5 billion parameter language model created by trinhkhng using the NuSLERP merge method. This model combines a base Qwen2-0.5B with a debiased version of Qwen2-0.5B, aiming to integrate their respective characteristics. It is designed for general language tasks where a compact model size and specific merged properties are beneficial, offering a 32768 token context length.

Loading preview...

Model Overview

The trinhkhng/nuslerp_Merged_Qwen2-0.5B_0.3 is a 0.5 billion parameter language model developed by trinhkhng. It was created using the NuSLERP merge method via mergekit, combining two distinct versions of the Qwen2-0.5B model.

Merge Details

This model is a blend of:

  • A base /kaggle/working/Qwen2-0.5B model, contributing 70% of the weight.
  • A /kaggle/working/debias_Qwen2-0.5B model, contributing 30% of the weight.

The NuSLERP method was configured with nuslerp_flatten: true and nuslerp_row_wise: false to achieve the desired integration of the constituent models. The tokenizer from the base Qwen2-0.5B model was utilized.

Key Characteristics

  • Architecture: Based on the Qwen2-0.5B family.
  • Parameter Count: 0.5 billion parameters, making it a compact model.
  • Context Length: Supports a context window of 32768 tokens.
  • Merge Method: Utilizes the NuSLERP technique to combine model weights, potentially integrating the strengths of both a standard and a debiased Qwen2-0.5B.

Potential Use Cases

This merged model is suitable for applications requiring a small, efficient language model that benefits from the combined characteristics of its base components. It could be particularly useful in scenarios where the debiasing efforts of one of the merged models are advantageous, while maintaining the general capabilities of the Qwen2-0.5B architecture.