jakeroxs/KAT-Coder-V2.5-Dev-35B-A3B-MTP-ABLITERATED

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:3Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 3, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

jakeroxs/KAT-Coder-V2.5-Dev-35B-A3B-MTP-ABLITERATED is a 35.51 billion parameter Qwen3.5 MoE model, with 3 billion active parameters and a native context of 262,144 tokens. Developed by jakeroxs, this model merges an abliterated variant of KAT-Coder V2.5 Dev with a Multi-Token Prediction (MTP) layer, featuring 1 MTP prediction layer. It is primarily intended for advanced coding workflows, agentic coding, and research into speculative decoding, offering reduced refusal behavior while retaining strong coding capabilities.

Loading preview...

KAT-Coder V2.5 Dev 35B-A3B MTP Abliterated Overview

This model, developed by jakeroxs, is a merged Hugging Face checkpoint based on the Qwen3.5 MoE architecture with approximately 35.51 billion parameters (3 billion active). It integrates a "Philadelphia Class abliteration" of the original KAT-Coder V2.5 Dev 35B-A3B with a compatible Multi-Token Prediction (MTP) layer from KAT-Coder V2.5 Dev, resulting in 1 MTP prediction layer.

Key Features and Modifications

  • Abliteration: Inherits modifications from the Philadelphia Class checkpoint, which aims to reduce refusal behavior by altering model representations without retraining the core model.
  • Multi-Token Prediction (MTP): Incorporates 19 MTP tensors, enabling potential advancements in token generation efficiency and speculative decoding.
  • Architecture: Utilizes a Qwen3.5 MoE architecture with 256 experts, using 8 experts per token, and boasts a substantial native context length of 262,144 tokens.
  • Lineage: A complex merge originating from KwaiPilot's KAT-Coder V2.5 Dev, further modified by KridgeDookie's abliterated variant, and enhanced with Myric's MTP head.

Intended Use Cases

This model is specifically provided for:

  • Further conversion or quantization, particularly for generating llama.cpp GGUF versions.
  • Experimentation with its transplanted MTP layer.
  • Local coding and agentic coding workflows, leveraging its underlying coding capabilities from KAT-Coder V2.5 Dev.
  • Research into speculative decoding techniques.

For direct llama.cpp usage, the associated GGUF repository is recommended: jakeroxs/KAT-Coder-V2.5-Dev-35B-A3B-MTP-ABLITERATED-GGUF.