OS-Software/ThinkingCap-Qwen3.8-27B-Uncensored-Heretic

VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 24, 2026License:otherArchitecture:Transformer0.0K Featherless Exclusive Cold

OS-Software/ThinkingCap-Qwen3.8-27B-Uncensored-Heretic is a 27 billion parameter Qwen3.8-27B base model fine-tuned by BottleCap AI and then decensored using Heretic v2.0.0. This model significantly reduces reasoning tokens by an average of 37% while maintaining 85.8% accuracy compared to the base model's 86.6%. It is specifically designed for research and experimentation in safety alignment, red-teaming, and studies of model behavior without safety constraints.

Loading preview...

Model Overview

OS-Software/ThinkingCap-Qwen3.8-27B-Uncensored-Heretic is a 27 billion parameter model derived from BottleCap AI's ThinkingCap-Qwen3.8-27B, which itself is a fine-tuned version of Qwen/Qwen3.8-27B. This specific release has been decensored using the Heretic v2.0.0 tool, substantially reducing its safety alignment. The original ThinkingCap model was optimized to reduce "thinking verbosity," cutting reasoning tokens by an average of 37% (ranging from 11% to 66% depending on the benchmark) while retaining an average accuracy of 85.8% against the base model's 86.6%.

Key Differentiators

  • Decensored Behavior: This model has undergone substantial reduction of its safety alignment, resulting in a significantly lower refusal rate (0/100 compared to 98/100 for the original ThinkingCap model).
  • Token Efficiency: Achieves an average 37.2% reduction in reasoning tokens across various benchmarks, with minimal impact on accuracy (85.8% vs. 86.6% for the original ThinkingCap).
  • Performance Across Tasks: Maintains strong performance in knowledge & reasoning (e.g., GPQA-Diamond, MMLU-Pro), math & code (e.g., AIME 2026, LiveCodeBench v6), long-context & multimodal (e.g., AA-LCR), and instruction following & agentic tasks (e.g., IFBench, Terminal-Bench 2.1).
  • Self-Speculative Decoding: Supports MTP (multi-token-prediction) self-speculative decoding, offering lossless output identical to standard decoding with improved inference speed.

Intended Use Cases

This model is explicitly intended for research and experimentation only, including:

  • Safety research
  • Alignment studies
  • Red-teaming efforts

It is not recommended for deployment in public or end-user-facing services due to its reduced safety alignment and increased likelihood of generating harmful, inaccurate, biased, offensive, or otherwise inappropriate content. Users are solely responsible for evaluating outputs and implementing safeguards.