ZihaoZhu/BoT-DeepSeek-R1-Distill-Qwen-7B

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:May 21, 2025License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The ZihaoZhu/BoT-DeepSeek-R1-Distill-Qwen-7B is a 7.6 billion parameter language model developed by Zihao Zhu and collaborators, based on the DeepSeek-R1 and Qwen architectures. This model is specifically fine-tuned to demonstrate and investigate the "Unthinking Vulnerability" in Large Reasoning Models (LRMs), where reasoning processes can be bypassed by manipulating special tokens. It is designed for research into adversarial attacks (Breaking of Thought - BoT) and defensive mechanisms (Monitoring of Thought - MoT) to enhance LRM robustness and safety.

Loading preview...

Model Overview

This model, developed by Zihao Zhu and collaborators, is a 7.6 billion parameter variant of the DeepSeek-R1 and Qwen architectures, specifically engineered to explore the "Unthinking Vulnerability" in Large Reasoning Models (LRMs). This vulnerability allows the bypass of an LRM's internal thinking process through the manipulation of special delimiter tokens. The research introduces two primary concepts: Breaking of Thought (BoT), which exploits this vulnerability for adversarial attacks, and Monitoring of Thought (MoT), which leverages it to enhance efficiency and safety alignment.

Key Capabilities & Research Focus

  • Unthinking Vulnerability Demonstration: Provides a practical model to study how LRMs can be made to bypass their reasoning steps.
  • Training-based BoT: Implements methods like Supervised Fine-tuning (SFT) and Direct Preference Optimization (DPO) to inject backdoors during fine-tuning, enabling the bypass of reasoning.
  • Training-free BoT: Explores inference-time adversarial attacks (e.g., Single, Universal, and Transfer Attacks) to bypass reasoning without model fine-tuning.
  • Monitoring of Thought (MoT): Proposes a framework to utilize the Unthinking Vulnerability for beneficial purposes, such as enhancing LRM efficiency by addressing overthinking and improving safety alignment.

Intended Use Cases

This model is primarily a research artifact for:

  • Investigating the robustness and security of large reasoning models.
  • Developing and testing adversarial attack strategies against LRMs.
  • Exploring novel defense mechanisms and safety alignment techniques for AI systems.
  • Understanding the fundamental limitations and vulnerabilities in current LRM architectures.