MassivDash/Qwen3-1.7B-heretic

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jan 8, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Warm

MassivDash/Qwen3-1.7B-heretic is a 1.7 billion parameter causal language model, a decensored version of Qwen/Qwen3-1.7B created using Heretic v1.1.0. It features a 32,768 token context length and is designed to reduce refusals compared to its original counterpart. This model retains Qwen3's advanced reasoning, instruction-following, and agent capabilities, with a unique ability to switch between thinking and non-thinking modes for diverse tasks.

Loading preview...

Model Overview

MassivDash/Qwen3-1.7B-heretic is a 1.7 billion parameter causal language model, derived from Qwen/Qwen3-1.7B and processed with Heretic v1.1.0 to create a decensored version. It maintains the original Qwen3's robust features while significantly reducing refusals, demonstrating 76 refusals out of 100 compared to the original's 91/100.

Key Capabilities & Features

  • Decensored Output: Modified to reduce content refusals, offering more direct responses.
  • Flexible Thinking Modes: Uniquely supports seamless switching between 'thinking mode' for complex logical reasoning, math, and coding, and 'non-thinking mode' for efficient, general-purpose dialogue. This can be controlled via enable_thinking parameter or /think and /no_think tags in prompts.
  • Enhanced Reasoning: Excels in mathematics, code generation, and commonsense logical reasoning, surpassing previous Qwen models.
  • Superior Human Preference Alignment: Strong in creative writing, role-playing, multi-turn dialogues, and instruction following.
  • Advanced Agent Capabilities: Designed for precise integration with external tools, achieving leading performance in complex agent-based tasks.
  • Multilingual Support: Supports over 100 languages and dialects with strong instruction following and translation abilities.
  • Extended Context Length: Features a substantial 32,768 token context window.

Use Cases & Best Practices

This model is suitable for applications requiring reduced content filtering and advanced reasoning, especially where dynamic control over the model's thought process is beneficial. For optimal performance, specific sampling parameters are recommended for thinking (Temperature=0.6, TopP=0.95) and non-thinking modes (Temperature=0.7, TopP=0.8). It is also advised to use an adequate output length (up to 38,912 tokens for complex problems) and standardize output formats for benchmarking, particularly for math and multiple-choice questions.