reaperdoesntknow/Dualmind-Qwen-1.7B-Thinking

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Mar 30, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Dualmind-Qwen-1.7B-Thinking is a 2 billion parameter Qwen3ForCausalLM model developed by Convergent Intelligence LLC: Research Division, fine-tuned using the DualMind SFT methodology on over 2.5 million tokens of Claude Opus 4.6 reasoning traces. This model specializes in extended deliberation and self-correction, absorbing the nuanced reasoning patterns of a frontier model, including backtracking and synthesis. It is built upon the DISC-refined Disctil-Qwen3-1.7B base model and is optimized for tasks requiring complex, multi-phase reasoning rather than simple pattern completion.

Loading preview...

Dualmind-Qwen-1.7B-Thinking Overview

Dualmind-Qwen-1.7B-Thinking is a 2 billion parameter Qwen3ForCausalLM model from Convergent Intelligence LLC: Research Division, specifically designed to emulate the sophisticated reasoning processes of Claude Opus 4.6. It was trained using the DualMind SFT methodology on a curated dataset of over 2.5 million tokens of Claude Opus 4.6 reasoning traces, focusing on extended deliberation and self-correction. This model distinguishes itself by learning the "shape of deliberation"—how a frontier model navigates uncertainty, backtracks, reconsiders, and synthesizes information.

Key Capabilities

  • Advanced Reasoning: Absorbs complex reasoning patterns, including multi-phase deliberation, self-correction, and nuanced uncertainty navigation.
  • Deliberative Structure: Produces outputs that reflect genuine thought processes, such as exploring, examining, and responding, rather than just pattern completion.
  • Robust Foundation: Built on the DISC-refined Disctil-Qwen3-1.7B base model, ensuring a strong structural foundation.
  • Extended Generation: Capable of generating long reasoning chains, supporting up to 40,960 tokens context length.

Good For

  • Complex Problem Solving: Ideal for applications requiring models to demonstrate deep, deliberative thought processes.
  • Cognitive Simulation: Useful for research into how smaller models can absorb and reproduce advanced cognitive loops from larger, more capable teachers.
  • Tasks Requiring Nuance: Excels in scenarios where simple, direct answers are insufficient and a more nuanced, self-correcting approach is beneficial.
  • Exploratory AI: Suitable for use cases where the model needs to explore multiple lines of thought before arriving at a conclusion.