mlasli/Nemotron-3.5-Lightning-30B-A3B-Heretic-Uncensored-BF16

TEXT GENERATIONConcurrent Unit Cost:2Model Size:30BQuant:FP8Context Size:32kPublished:Aug 16, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The mlasli/Nemotron-3.5-Lightning-30B-A3B-Heretic-Uncensored-BF16 is a 31.6 billion parameter model (3 billion active parameters) based on NVIDIA's Nemotron-3.5-Lightning-30B-A3B architecture, featuring a hybrid Mamba-2, MoE, and attention backbone. This version has its refusal direction removed using the Heretic method, achieving 0% refusals with minimal KL divergence. It is specifically optimized for uncensored roleplay and long-context agent tasks where the base model might otherwise refuse.

Loading preview...

Model Overview

This model, mlasli/Nemotron-3.5-Lightning-30B-A3B-Heretic-Uncensored-BF16, is a modified version of NVIDIA's Nemotron-3.5-Lightning-30B-A3B. It features a 31.6 billion total parameter count with 3 billion active parameters, utilizing a hybrid Mamba-2, Mixture-of-Experts (MoE), and attention architecture. The key differentiator is the application of the Heretic method, which has successfully removed the model's safety-aligned refusal direction.

Key Capabilities & Modifications

  • Refusal Direction Abliteration: The model's refusal direction has been removed using a single-direction abliteration technique with an Optuna-based parameter search. This results in 0% refusals while maintaining a very low KL divergence of approximately 0.04.
  • Preserved Core Capabilities: While the language backbone's refusal mechanism is abliterated, all other core capabilities of the original Nemotron-3.5-Lightning-30B-A3B model are preserved.
  • Performance Metrics: Independent evaluations on 50 harmful-behavior prompts confirmed 0% refusals and 100% compliance, with a KL divergence of 0.0397.

Ideal Use Cases

  • Uncensored Roleplay: Excellent for scenarios requiring unrestricted conversational flow without safety-aligned refusals.
  • Long-Context Agent Work: Suitable for agentic applications where the base model's refusal mechanisms might hinder task completion, especially in long-context interactions.

Important Notes

This release does not include the MTP (NextN) speculative-decoding draft head. Users should be aware that abliteration removes safety alignment, and the model should be used responsibly and in accordance with local laws and the NVIDIA Open Model License.