Disya/Mistral-qwq-12b-merge

Hugging Face
TEXT GENERATIONPricing:Input $0.87 / Cached $0.2 / Output $0.99Concurrent Unit Cost:1Model Size:12BQuant:FP8Context Size:32kTool Calling:SupportedPublished:May 12, 2025Architecture:Transformer0.0K Featherless Exclusive Warm

Disya/Mistral-qwq-12b-merge is a 12 billion parameter language model created by Disya, merged using the DARE TIES method. It combines several pre-trained models, including BeaverAI/MN-2407-DSK-QwQify-v0.1-12B, CreitinGameplays/Mistral-Nemo-12B-R1-v0.2, Nitral-AI/Mag-Mell-Reasoner-12B, Dans-DiscountModels/12b-mn-dans-reasoning-test-5, and Delta-Vector/Francois-PE-V2-Huali-12B. This merge aims to leverage the strengths of its constituent models, offering a unique blend of capabilities for various language generation tasks.

Loading preview...

Mistral-qwq-12b-merge Overview

Disya/Mistral-qwq-12b-merge is a 12 billion parameter language model resulting from a sophisticated merge of multiple pre-trained models. This model was constructed using the DARE TIES merge method, which combines the weights of several base models to create a new, potentially more capable model. The primary base model for this merge was BeaverAI/MN-2407-DSK-QwQify-v0.1-12B.

Key Capabilities

  • Blended Expertise: Integrates knowledge and capabilities from five distinct 12B parameter models, including those focused on reasoning and general language understanding.
  • DARE TIES Merge Method: Utilizes a specific merging technique designed to combine models effectively, potentially enhancing overall performance.
  • Customizable Foundation: Built upon a diverse set of foundational models, suggesting a broad range of potential applications.

Good For

  • Exploratory Research: Ideal for researchers and developers interested in experimenting with merged models and their emergent properties.
  • Diverse Language Tasks: Suitable for general-purpose language generation, understanding, and reasoning tasks, benefiting from the combined strengths of its components.
  • Custom Model Development: Provides a strong base for further fine-tuning or specialized applications where a blend of different model characteristics is desired.