saidutta69/Meta-Llama-3.1-8B-Instruct-heretic

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Jul 21, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The saidutta69/Meta-Llama-3.1-8B-Instruct-heretic is a decensored variant of Meta-Llama-3.1-8B-Instruct, created by saidutta69 using the Heretic v1.4.0 abliteration method. This model suppresses refusal behavior via targeted weight edits, leaving the base model's knowledge and instruction-following capabilities largely intact. It is designed for developers seeking a Llama 3.1 8B model that provides direct answers without built-in refusal guardrails, suitable for deployment on 8-12 GB GPUs or consumer hardware with GGUF quantizations.

Loading preview...

Overview

saidutta69/Meta-Llama-3.1-8B-Instruct-heretic is a modified version of the Meta-Llama-3.1-8B-Instruct model, developed by saidutta69. Its primary distinction is the removal of refusal behaviors through a process called "abliteration" using the Heretic v1.4.0 tool. This method involves targeted weight edits to the attention output and MLP down-projections, which suppresses refusal without fine-tuning, thus preserving the base model's original knowledge and instruction-following abilities.

Key Capabilities & Features

  • Decensored Responses: Deliberately suppresses refusal behavior, allowing the model to comply with requests that the base model would typically refuse.
  • Preserved Base Model Quality: Maintains the core knowledge and instruction-following capabilities of the original Meta-Llama-3.1-8B-Instruct.
  • Abliteration Method: Utilizes a unique technique that edits specific weight directions responsible for refusal, avoiding the potential coherence degradation associated with fine-tuning.
  • Hardware Compatibility: Designed to run on 8-12 GB GPUs and supports GGUF quantizations (Q4_K_M, Q5_K_M, Q6_K, Q8_0) for consumer hardware.

When to Use This Model

This model is ideal for developers who require a Llama 3.1 8B model that provides direct answers without built-in safety guardrails or refusal mechanisms. It is suitable for use cases where the user is responsible for content moderation and desires a model that will attempt to answer all prompts. It is not a capability upgrade over the base model but rather a modification of its response behavior.