TheZeez/gemma-4-e4b-creative-DFT-exp
VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:May 19, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold
TheZeez/gemma-4-e4b-creative-DFT-exp is a 7.9 billion parameter Gemma 4 E4B model, custom-trained by TheZeez for creative writing tasks. It utilizes a novel Distribution Fine-Tuning (DFT) method to mathematically penalize and reduce repetitive AI-generated phrasing. This model is specifically designed to produce more diverse and less predictable text, making it suitable for generating unique creative content.
Loading preview...
Model Overview
This model, TheZeez/gemma-4-e4b-creative-DFT-exp, is a custom-trained version of Google's Gemma 4 E4B, featuring 7.9 billion parameters and a 32768-token context length. Its primary distinction lies in its training methodology, which employs a unique Distribution Fine-Tuning (DFT) approach.
Key Capabilities & Innovations
- Eliminates Repetitive AI Slop: Unlike standard SFT and RLHF, this model was trained with a macro-statistical loss penalty. It actively penalizes the overuse of common AI-frequent vocabulary (e.g., "whisper", "shiver") by calculating the Mean Squared Error (MSE) between the model's batch-level vocabulary distribution and a human target distribution.
- Enhanced Creative Writing: The DFT method aims to produce more diverse and less predictable text, making it particularly effective for creative writing applications.
- Training Details: The model underwent approximately 1.3 epochs of training with an effective batch size of 96, utilizing a dataset composed of multiturn roleplaying and creative writing examples.
Use Cases
- Creative Content Generation: Ideal for generating unique stories, dialogues, and other creative texts that require varied phrasing and avoid generic AI patterns.
- Roleplaying Scenarios: Suitable for developing dynamic and less predictable responses in roleplaying applications due to its training on multiturn roleplaying datasets.