trinhkhng/nearswap_Merged_Qwen2-0.5B_0.0

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 6, 2026Architecture:Transformer Featherless Exclusive Cold

The trinhkhng/nearswap_Merged_Qwen2-0.5B_0.0 is a 0.5 billion parameter language model, merged using the NearSwap method with a Qwen2-0.5B base. This model specifically incorporates a debiased version of Qwen2-0.5B, aiming to provide a foundational model with reduced biases. It is suitable for applications requiring a compact, debiased language model, particularly where the base Qwen2 architecture is preferred.

Loading preview...

Model Overview

The trinhkhng/nearswap_Merged_Qwen2-0.5B_0.0 is a 0.5 billion parameter language model created by trinhkhng. It is a product of a model merging process, specifically utilizing the NearSwap method. The base model for this merge was /kaggle/working/Qwen2-0.5B, and it incorporated /kaggle/working/debias_Qwen2-0.5B as a component.

Key Characteristics

  • Merge Method: Employs the NearSwap technique, which is designed to combine pre-trained language models effectively.
  • Base Architecture: Built upon the Qwen2-0.5B architecture, providing a compact and efficient foundation.
  • Debiased Component: Integrates a debiased version of Qwen2-0.5B, suggesting an effort to mitigate inherent biases present in the original model.
  • Parameter Count: Features 0.5 billion parameters, making it a relatively small and fast model suitable for resource-constrained environments or specific tasks.
  • Context Length: Supports a context length of 32768 tokens, allowing it to process substantial amounts of input text.

Potential Use Cases

  • Bias Mitigation Research: Useful for exploring the effects of debiasing techniques on language models.
  • Resource-Constrained Applications: Its small size makes it suitable for deployment on devices with limited computational power.
  • Foundational NLP Tasks: Can serve as a base for various natural language processing tasks where a compact, debiased model is beneficial.