SubMaroon/Gemma-4-26B-A4B-StyleTune-QK-Heretic

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:2Model Size:26BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 13, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

SubMaroon/Gemma-4-26B-A4B-StyleTune-QK-Heretic is a 26 billion parameter Gemma-4-26B-A4B variant with a 32768 token context length, featuring specific architectural edits. It incorporates a replaced lm_head and attention routing interpolated towards a roleplay finetune, making it suitable as a base for further finetuning and merges. This model is designed to improve prompt parsing and contextual understanding, particularly in narrative-driven scenarios.

Loading preview...

Overview

SubMaroon/Gemma-4-26B-A4B-StyleTune-QK-Heretic is a specialized 26 billion parameter Gemma-4-26B-A4B model, distinguished by three key modifications: ablated refusal directions, a replaced lm_head from Gryphe/Gemma-4-26B-A4B-StyleTune-V2, and attention routing interpolated towards a roleplay finetune. This model is released primarily as a robust starting point for developers looking to create further finetunes and merges.

Key Architectural Edits

  • Body: Derived from coder3101/gemma-4-26B-A4B-it-heretic (Heretic ARA, layers 10-30).
  • lm_head: Replaced with one from Gryphe/Gemma-4-26B-A4B-StyleTune-V2.
  • q_proj, k_proj: Interpolated using a task vector derived from Pantheon-Reasoning-1.1-V2 to enhance attention routing.
  • MoE experts, router, embeddings, MLP, and the vision tower remain identical to the original abliterated body.

Performance Characteristics

  • Improved Prompt Parsing: The model demonstrates enhanced ability to act on specific prompt situations, including participants and actions, leading to more contextually relevant generations.
  • Lexical Diversity: A slight, potentially negligible, drop in lexical diversity was observed.
  • Length: Mean prose length and word length remain largely unchanged compared to the pre-QK build.

Intended Use

This model is specifically designed as a starting point for further finetunes and merges, offering a modified base with improved contextual understanding and prompt adherence, particularly beneficial for roleplay or narrative generation tasks.