blackbook-lm/DeepSeek-R1-Distill-Qwen-32B-heretic

TEXT GENERATIONConcurrent Unit Cost:2Model Size:32.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Mar 1, 2026License:mitArchitecture:Transformer0.0K Open Weights Featherless Exclusive Cold

blackbook-lm/DeepSeek-R1-Distill-Qwen-32B-heretic is a 32.8 billion parameter language model, a decensored version of DeepSeek-R1-Distill-Qwen-32B. This model is part of the DeepSeek-R1-Distill series, which distills reasoning patterns from the larger DeepSeek-R1 model into smaller, dense architectures. It is specifically modified to reduce refusals, making it suitable for applications requiring less restrictive content generation while maintaining strong reasoning capabilities in math, code, and general English tasks.

Loading preview...

Model Overview

This model, blackbook-lm/DeepSeek-R1-Distill-Qwen-32B-heretic, is a 32.8 billion parameter language model derived from deepseek-ai/DeepSeek-R1-Distill-Qwen-32B. It has been processed using the Heretic v1.2.0 tool to create a "decensored" version, significantly reducing refusal rates compared to its original counterpart (3/100 vs. 56/100 refusals).

Key Capabilities & Distinguishing Features

  • Decensored Output: Modified to produce fewer refusals, offering more direct and less restricted responses.
  • Reasoning Distillation: Benefits from reasoning patterns distilled from the larger DeepSeek-R1 model, which was trained using large-scale reinforcement learning without initial supervised fine-tuning.
  • Strong Performance: The base DeepSeek-R1-Distill-Qwen-32B model demonstrates competitive performance across various benchmarks, including math (AIME 2024 pass@1: 72.6%, MATH-500 pass@1: 94.3%), code (LiveCodeBench pass@1: 57.2%), and general reasoning (GPQA Diamond pass@1: 62.1%).
  • Qwen2.5 Base: Built upon the Qwen2.5-32B architecture, leveraging its foundational capabilities.

When to Use This Model

This model is particularly well-suited for use cases where a less restrictive and more direct response generation is desired, especially in applications requiring strong reasoning in mathematical, coding, and general knowledge domains. Its reduced refusal rate makes it a candidate for tasks where the original model might have been overly cautious or restrictive.