orlandorubino/Qwen3-1.7B-heretic

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 29, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

orlandorubino/Qwen3-1.7B-heretic is a 2 billion parameter Qwen3-based causal language model, developed by orlandorubino, specifically modified to reduce refusal behaviors. This model utilizes the Heretic method to ablate censorship, resulting in significantly fewer refusals while maintaining a low KL divergence from the original Qwen3-1.7B. It is optimized for use cases requiring a less restricted language model, offering a 32768 token context length.

Loading preview...

Overview

orlandorubino/Qwen3-1.7B-heretic is a modified version of the Qwen/Qwen3-1.7B model, specifically engineered to reduce its refusal tendencies. This "heretic" variant was created using the Heretic tool, which identifies and nullifies the "rejection direction" in the model's activation space through orthogonalization.

Key Modifications & Results

  • Censorship Ablation: The Heretic process was applied to remove refusal safeguards, comparing harmless prompts with those that typically trigger rejections. Optuna was used to find parameters that minimize refusals while preserving the original model's capabilities, indicated by a low KL divergence.
  • Efficiency: The process involved 25 Optuna trials, completed in approximately 45 minutes on an 8GB GPU (RTX 2060 Super) using BF16 precision.

Performance Metrics

Metric Original Heretic
Refusals (bad prompts) 92/100 3/100
KL Divergence 0 0.0566

Usage

The model can be loaded using the transformers library, with options for 4-bit quantization for smaller GPUs. Users should be aware that this model has reduced refusal safeguards and are responsible for its use in accordance with the law and the base model's Apache 2.0 license.