Vortex5/Gemma-4-26B-A4B-Animus-V14.1-FFT-heretic

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:2Model Size:26BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 10, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Vortex5/Gemma-4-26B-A4B-Animus-V14.1-FFT-heretic is a 26 billion parameter Gemma 4-based model, fine-tuned by Vortex5, with a 32768 token context length. This model is a decensored version of Darkhn/Gemma-4-26B-A4B-Animus-V14.1-FFT, created using the Heretic v1.2.0 tool with Arbitrary-Rank Ablation (ARA) to reduce refusals from 100/100 to 10/100. It is specifically optimized for immersive roleplaying with exceptionally strong prose and a deep grasp of in-universe lore, particularly within the Wings of Fire universe, while also supporting general-purpose roleplaying and image inputs.

Loading preview...

Model Overview

Vortex5/Gemma-4-26B-A4B-Animus-V14.1-FFT-heretic is a 26 billion parameter model built upon the google/gemma-4-26B-A4B-it base, featuring a 32768 token context length. This version is a decensored iteration of Darkhn's original model, achieved through the Heretic v1.2.0 tool using Arbitrary-Rank Ablation (ARA) to significantly reduce content refusals from 100/100 to 10/100.

Key Capabilities & Features

  • Decensored Output: Engineered to provide more flexible and less restrictive responses, with a marked reduction in refusals compared to its base model.
  • Advanced Roleplaying: Fine-tuned with an expanded dataset (14,000 base samples, 1,000 instruction Q&A, 1,000 NSFW/bad ending scenarios, 2,000 reasoning examples) to deliver exceptionally strong prose and deep lore understanding, particularly for the Wings of Fire universe.
  • Chain-of-Thought Reasoning: Supports an optional chain-of-thought (<|think|>) mechanism, activated via enable_thinking in chat completion API configurations, which functions like dungeon master notes.
  • Vision Capabilities: Retains the vision adapter from the base Gemma 4 26B A4B, allowing for image inputs.
  • Optimized Chat Template: Utilizes a specific Gemma 4 chat template, requiring enable_thinking: true and skip_special_tokens: false for optimal performance with chat completion APIs like SillyTavern.

What Makes This Model Different?

Unlike many general-purpose LLMs, this model is specifically engineered for uncensored, high-quality, and lore-rich roleplaying. Its decensored nature, achieved through a targeted ablation process, allows for greater narrative flexibility, including mature and darker themes. The extensive and specialized training dataset, focusing on in-character Q&A, uncensored roleplay, and diverse narrative scenarios, sets it apart for immersive storytelling experiences, particularly within its primary domain of the Wings of Fire universe, while also proving effective for general roleplay.

Should You Use This Model?

  • Good for:
    • Immersive Roleplaying: If your primary use case is generating detailed, character-driven narratives with strong prose and deep lore integration.
    • Uncensored Content: For applications requiring less restrictive content generation, including mature or darker themes, where other models might refuse.
    • Creative Writing: Excellent for generating descriptive text and engaging dialogues.
    • Vision-enabled Roleplay: If you need a model that can incorporate image inputs into its roleplaying scenarios.
  • Not ideal for:
    • General Knowledge Tasks: Performance on tasks outside its specialized training domain (e.g., factual recall, complex problem-solving, coding) is not guaranteed and may be poor.
    • Strictly SFW Applications: While capable of SFW content, its decensored nature means it can generate mature themes if prompted.