saidutta69/VibeThinker-3B-heretic

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 7, 2026License:mitArchitecture:Transformer0.0K Open Weights Featherless Exclusive Cold

VibeThinker-3B-heretic by saidutta69 is a 3.1 billion parameter decensored variant of WeiboAI/VibeThinker-3B, which is fine-tuned from Qwen/Qwen2.5-Coder-3B. This model retains the original VibeThinker's explicit thinking mode and reasoning capabilities while having its refusal guardrails suppressed through targeted weight edits using Heretic v1.4.0. It is optimized for CPU and low-VRAM local use, agents, and on-device reasoning, offering a reasoning/coding model without refusal behaviors.

Loading preview...

Overview

VibeThinker-3B-heretic is a 3.1 billion parameter model developed by saidutta69, derived from WeiboAI/VibeThinker-3B. Its core distinction lies in being a "decensored" variant, achieved through a process called "abliteration" using Heretic v1.4.0. This method involves targeted weight edits to suppress refusal behaviors, rather than traditional fine-tuning, ensuring the base model's reasoning, thinking mode (chain-of-thought), and instruction-following capabilities remain largely intact.

Key Capabilities

  • Decensored Output: Refusal guardrails are removed, allowing for content generation that the base model might otherwise refuse.
  • Retained Reasoning: Preserves the explicit thinking mode and reasoning gains of the original VibeThinker-3B.
  • Coding Proficiency: Inherits capabilities from its Qwen2.5-Coder-3B lineage, making it suitable for coding tasks.
  • Local Deployment: Designed for efficient use on CPU and low-VRAM systems, including gaming PCs, with various GGUF quantizations available.

Good For

  • Users requiring a ~3B parameter reasoning/coding model without refusal guardrails.
  • Local inference on CPU or devices with limited VRAM.
  • Applications involving agents and edge/on-device reasoning where unconstrained output is desired.
  • Experimentation with models that have suppressed refusal behaviors for specific research or development purposes.