NostraEmpire/mirror-qwen3-1.7b

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 31, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Qwen3-1.7B is a 1.7 billion parameter causal language model from the Qwen series, developed by Qwen. It features a unique ability to seamlessly switch between a 'thinking mode' for complex reasoning (math, code) and a 'non-thinking mode' for general dialogue, enhancing performance across diverse tasks. This model excels in reasoning, instruction-following, agent capabilities, and multilingual support for over 100 languages, making it suitable for applications requiring adaptable intelligence.

Loading preview...

Qwen3-1.7B Model Overview

Qwen3-1.7B is a 1.7 billion parameter causal language model developed by Qwen, part of the latest Qwen series. It introduces a novel feature allowing seamless switching between a 'thinking mode' for complex logical reasoning, mathematics, and coding, and a 'non-thinking mode' for efficient, general-purpose dialogue. This adaptability ensures optimized performance across various scenarios.

Key Capabilities and Differentiators

  • Adaptive Reasoning: Uniquely supports dynamic mode switching, significantly enhancing performance in mathematical problem-solving, code generation, and commonsense reasoning compared to previous Qwen models.
  • Superior Alignment: Demonstrates strong human preference alignment, excelling in creative writing, role-playing, multi-turn conversations, and instruction following for a more natural user experience.
  • Advanced Agentic Abilities: Offers robust agent capabilities, integrating precisely with external tools in both thinking and non-thinking modes, achieving leading performance in complex agent-based tasks among open-source models.
  • Extensive Multilingual Support: Supports over 100 languages and dialects, providing strong capabilities for multilingual instruction following and translation.
  • Context Length: Features a substantial context length of 32,768 tokens.

Recommended Use Cases

This model is particularly well-suited for developers building applications that require:

  • Dynamic Task Handling: Scenarios where the model needs to switch between deep analytical reasoning and quick, general responses.
  • Complex Problem Solving: Applications involving mathematics, coding, or intricate logical reasoning.
  • Engaging Conversational AI: Chatbots or interactive agents that require superior human preference alignment, creative writing, or role-playing.
  • Tool Integration: Agent-based systems that leverage external tools for complex workflows.
  • Multilingual Applications: Projects requiring robust understanding and generation across a wide array of languages.