SubMaroon/Kanimus-26B-A4B-FFT-heretic

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:2Model Size:26BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 27, 2026License:gemmaArchitecture:Transformer0.0K Featherless Exclusive Cold

SubMaroon/Kanimus-26B-A4B-FFT-heretic is an experimental 26 billion parameter language model based on an abliterated Gemma-4-26B-A4B-Animus-V14.1-FFT base, designed for dark roleplay scenarios. It incorporates a task-arithmetic delta from a Claude Opus distillation into its attention Q and K projections, alongside a baked-in stylistic LoRA. This model aims to modify stylistic output for roleplay while exhibiting altered instruction following due to its experimental merge methodology, supporting a 32768 token context length.

Loading preview...

Kanimus-26B-A4B-FFT-heretic: Experimental Dark Roleplay Merge

This model, developed by SubMaroon, is an experimental 26 billion parameter merge built upon the Vortex5/Gemma-4-26B-A4B-Animus-V14.1-FFT-heretic base. It is specifically designed for dark roleplay applications, incorporating two key modifications to influence its output style and behavior.

Key Modifications and Architecture

The model integrates a task-arithmetic delta into its q_proj and k_proj attention projections. This delta is derived from a Claude Opus distillation (TeichAI/gemma-4-26B-A4B-it-Claude-Opus-Distill-v2) using unsloth/gemma-4-26B-A4B-it as the parent. Additionally, a stylistic LoRA (SubMaroon/Dark-Goetia-26B-A4B-LoRA-v4) is baked into 205 modules, including attention projections and shared dense MLPs, at a scaled factor of 0.40.

Behavioral Characteristics

While no quantitative evaluations were performed, qualitative observations suggest:

  • Instruction following may be degraded compared to the base model, particularly in adhering to specific scene settings or user-provided facts.
  • Prose style remains generally close to the base, with some divergence in longer outputs.
  • Potential for referent drift (English) and token artifacts (Russian) has been noted in earlier testing, though not consistently reproduced.

Use Cases and Considerations

This model is intended for developers exploring experimental merges and those specifically targeting dark roleplay content generation. Users should be aware of potential deviations in instruction adherence. It supports a 32768 token context length and is available in BF16 and GGUF (Q4_K_M) formats. Sampling recommendations include a temperature of 0.7-0.8 and min_p of 0.05-0.1.