trinhkhng/linear_Merged_Qwen2-0.5B_0.2

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 6, 2026Architecture:Transformer Featherless Exclusive Cold

The trinhkhng/linear_Merged_Qwen2-0.5B_0.2 is a 0.5 billion parameter language model created by trinhkhng, merged from Qwen2-0.5B and a debiased variant using the Linear method. This model combines the base capabilities of Qwen2-0.5B with specific adjustments from the debiased model. It is suitable for applications requiring a compact, efficient language model with a 32768-token context length.

Loading preview...

Model Overview

The trinhkhng/linear_Merged_Qwen2-0.5B_0.2 is a 0.5 billion parameter language model developed by trinhkhng. It was constructed using the Linear merge method via mergekit, combining two distinct base models: /kaggle/working/Qwen2-0.5B and /kaggle/working/debias_Qwen2-0.5B.

Merge Details

This model is a composite, with the base Qwen2-0.5B contributing 80% of the weight and the debiased variant contributing 20%. This specific weighting suggests an intent to retain the core characteristics of the original Qwen2-0.5B while integrating improvements or modifications from the debiased version. The model maintains a substantial context length of 32768 tokens.

Key Characteristics

  • Architecture: Based on the Qwen2-0.5B family.
  • Parameter Count: 0.5 billion parameters, making it a relatively compact model.
  • Context Length: Supports a 32768-token context, allowing for processing longer inputs.
  • Merge Method: Utilizes the Linear merge technique to blend model weights.

Potential Use Cases

This model is well-suited for scenarios where a smaller, efficient language model is required, potentially benefiting from the debiasing efforts. It can be considered for:

  • Resource-constrained environments.
  • Tasks requiring a balance between performance and computational cost.
  • Applications where the specific characteristics introduced by the debiased component are advantageous.