SubMaroon/Boulesis-26B-A4B

Hugging Face
VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:2Model Size:26BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 4, 2026License:gemmaArchitecture:Transformer0.0K Featherless Exclusive Warm

SubMaroon/Boulesis-26B-A4B is a 26 billion parameter composite Gemma 4 model, developed by SubMaroon, specifically engineered for advanced roleplay scenarios. It utilizes QK task arithmetic and fused LoRA to enhance narrative progression, character understanding, and context retention. This model excels at generating dynamic and decisive prose while maintaining core intelligence, making it ideal for complex interactive storytelling.

Loading preview...

Boulesis-26B-A4B: Advanced Roleplay Model

Boulesis-26B-A4B is a 26 billion parameter model developed by SubMaroon, built upon the Gemma 4 architecture. It is a composite model that leverages QK task arithmetic and fused LoRA to significantly enhance its roleplaying capabilities. The primary goal in its development was to create a model that moves beyond passive reactivity, actively advancing narratives and demonstrating a deeper understanding of character context.

Key Capabilities & Differentiators

  • Enhanced Narrative Progression: Designed to generate more decisive and dynamic prose, actively driving the story forward rather than just mirroring user input.
  • Superior Context Retention: Sharpened attention to context allows the model to organically integrate lore and facts from character cards into roleplay.
  • Zero Catastrophic Forgetting: Retains the base intelligence of its Gemma 4 foundation while specializing in roleplay.
  • Unique Merging Strategy: Unlike typical merges, Boulesis edits specific parts of the attention mechanism (q_proj, k_proj) and incorporates an lm_head from Gryphe/Gemma-4-26B-A4B-StyleTune-V2 to achieve its specialized behavior.

Version Differences

  • v1: Focuses on dynamic plots and action-oriented scenarios.
  • v2: Prioritizes concise writing while maintaining logical coherence.
  • v2.1: Offers the best grasp of characters and excellent memory for complex, multi-character roleplay sessions.

Recommended Usage

For optimal performance, users are advised to enable "Thinking" mode and use specific inference parameters: Temperature 1.0, Top-K 64, Repetition Penalty 1.05-1.1, and Top-P 0.95.