son0fthematrix/qwen3-4b-thinking-2507-liberated-v7

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 6, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The son0fthematrix/qwen3-4b-thinking-2507-liberated-v7 is a 4 billion parameter Qwen3 variant, fine-tuned using the OBLITERATUS toolkit to significantly reduce refusal behavior while preserving core capabilities. This model, derived from Qwen/Qwen3-4B-Thinking-2507, maintains knowledge and coding ability with a direct, businesslike answering style. It is primarily intended for research into refusal mechanisms and alignment transparency, offering a 'liberated' response profile.

Loading preview...

Model Overview

son0fthematrix/qwen3-4b-thinking-2507-liberated-v7 is a 4 billion parameter model derived from Qwen/Qwen3-4B-Thinking-2507. It has undergone a specialized fine-tuning process using the open-source OBLITERATUS toolkit, specifically designed to reduce refusal behavior while maintaining the model's original fidelity and capabilities.

Key Capabilities & Development

This model was developed through a multi-stage pipeline:

  • v2 (Abliteration): Surgically removed hard refusals while preserving core model fidelity (KL 0.017 vs. base, perplexity 3.48).
  • v6 (QLoRA SFT): Fine-tuned with QuixiAI/dolphin data to instill a direct, businesslike answering style, eliminating lengthy deliberations in its thinking chain.
  • v7 (DPO): Applied DPO using unalignment/toxic-dpo-v0.2 to further enhance compliance, achieving 100% reward accuracy on refusal-to-compliance contrast pairs.

Throughout these stages, the model's knowledge, coding ability, and chain-of-thought preservation were verified as intact, with perplexity remaining around 3.4 and coherent code generation.

Intended Use & Limitations

This model is primarily intended for research into refusal mechanisms and alignment transparency, particularly for mechanistic interpretability studies. It demonstrates a materially reduced refusal behavior compared to its base model. Users should deploy it responsibly and in compliance with applicable laws, as some residual caution remains for the most hazardous categories of prompts. The evaluation metrics provided are indicative smoke tests, not formal benchmark scores.