Blackroot/Ammeg-26B-A4
VISIONConcurrent Unit Cost:2Model Size:26BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 19, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold
Blackroot/Ammeg-26B-A4 is an experimental finetune of the Gemma-26B-4A model, developed by Blackroot. This model utilizes a novel quality-preserving obliteration method followed by retraining with an evolutionary strategy. It is specifically designed for story generation, leveraging a non-differentiable objective that exploits the Zipf structure of language for long-term dependencies. The tuning process involved low-rank updates on a per-sample basis, resulting in a model optimized for creative narrative tasks.
Loading preview...
Ammeg-26B-A4: An Experimental Finetune
Blackroot/Ammeg-26B-A4 is an experimental finetune based on the Gemma-26B-4A architecture. This model introduces a unique two-part tuning process:
Key Capabilities & Innovations
- Novel Obliteration Method: Employs a quality-preserving obliteration technique, loosely based on 'heretic', to prepare the base model.
- Evolutionary Strategy Retraining: The obliterated model is then retrained using an evolutionary strategy, drawing inspiration from recent research (e.g., arXiv:2511.16652).
- Story Generation Focus: Data for finetuning primarily consists of books, with the obliterated model captioning stories to generate prompts like "Generate a story with a protagonist named Alice...".
- Unique Tuning Method: Utilizes low-rank (Rank 1 LORA) updates per sample into a full-precision buffer with stochastic rounding, aiming for noise cancellation and full-rank tuning in a Gaussian landscape.
- Non-Differentiable Objective: Incorporates a complex non-differentiable objective that exploits the Zipf structure of language as a proxy for long-term dependencies in narratives.
Important Considerations
- Safeguard Removal: The specific tuning method has likely removed the model's original safeguards, including tendencies to soft-refuse or redirect, meaning typical safety features may not be present.
- Hardware & Training: The training was conducted on powerful CPUs over approximately 280 hours, highlighting a slow but cost-effective approach.
Good For
- Experimental research into novel finetuning methodologies.
- Creative story generation and narrative tasks.
- Exploring models with unconventional training paradigms.