ApolloRaines/Qwen2.5-Coder-32B-Instruct-Jbliterated

TEXT GENERATIONConcurrent Unit Cost:2Model Size:32.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 14, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

ApolloRaines/Qwen2.5-Coder-32B-Instruct-Jbliterated is a 32 billion parameter instruction-tuned causal language model, based on the Qwen2.5-Coder architecture. Developed by Apollo Raines, this model has undergone "Jbliteration," a mechanistically-targeted process that removes refusal behavior while preserving the model's original personality, humor, and creative expression. It is designed as a drop-in replacement for the base Qwen2.5-Coder-32B-Instruct, excelling in code generation and creative tasks without censorship. This model is part of the B² (B-Squared) architecture research, focusing on structural innovation for enhanced performance.

Loading preview...

Model Overview

ApolloRaines/Qwen2.5-Coder-32B-Instruct-Jbliterated is a specialized version of the Qwen2.5-Coder-32B-Instruct model, developed by Apollo Raines. Its primary distinction is the application of Jbliteration, a novel method designed to surgically remove refusal behavior without compromising the model's inherent personality, humor, or creative voice. This contrasts with standard abliteration techniques that often result in a duller, less expressive model.

Key Innovations

  • Jbliteration: This process uses the Jacobian Lens to identify and remove only the causal components of refusal directions, ensuring that the model retains its original expressive qualities. It involves concept mining, Jacobian Lens extraction, restricted projection, and norm-preserving application across linear layers.
  • B² Architecture Project: This model is a component of the broader B² (B-Squared) architecture research, which aims to create AI systems that achieve high performance through structural innovation rather than brute-force scaling. Jbliteration plays a crucial role by providing cleaner foundation weights for this architecture.

Capabilities & Usage

This model functions as a direct replacement for the original Qwen2.5-Coder-32B-Instruct, maintaining the same architecture, tokenizer, and context length. It is particularly suited for tasks requiring uncensored output, such as creative writing, security research, and academic study, where preserving the model's full expressive range is critical. The model is available in BF16, Q8_0 GGUF, and Q4_K_M GGUF formats, supporting various deployment scenarios from GPU inference to CPU+GPU offload.

Responsible Use

Users are cautioned that all refusal behavior has been removed, making them solely responsible for the model's output. It is intended for legitimate use cases and research, with a strong emphasis on ethical and legal compliance.