sfsx/Z-Image-Engineer-V6
Z-Image-Engineer V6 by sfsx is a 4 billion parameter Qwen text encoder fine-tuned for Z-Image workflows. Optimized using SMART DoRA training, it excels at transforming minimal seed prompts into rich, structured visual narratives by adding explicit scene composition, lighting, and depth. This model serves as both a local prompt enhancement tool and a direct text encoder replacement for Z-Image, improving image generation quality.
Loading preview...
Overview
Z-Image-Engineer V6 is a 4 billion parameter Qwen text encoder, fine-tuned by BennyDaBall using the SMART DoRA training system. It is designed to enhance and structure image prompts for Z-Image workflows, transforming simple concepts into detailed visual narratives. The model can function as a local prompt enhancer or directly replace the stock Z-Image text encoder, offering dual-role performance.
Key Capabilities
- Prompt Enhancement: Rewrites minimal seed prompts into rich, highly structured visual descriptions, adding elements like scene composition, lighting, and material textures.
- Text Encoder Replacement: Can be used as a direct swap for the
Tongyi-MAI/Z-Image-Turbotext encoder to generate different conditioning from the same seed prompt. - Hybrid Mode: Allows for prompt rewriting and subsequent encoding using V6, enabling the model to both craft the scene and drive the image generation.
- Private Local Workflow: Built for local environments like LM Studio, ComfyUI, and
llama.cpp, ensuring no API logs or external telemetry. - Advanced Training: Utilizes Weight-Decomposed Low-Rank Adaptation (DoRA) with SMART regularizers (Entropic, Holographic, Topological, Manifold) to prevent repetitive outputs and enforce structured logic.
Good For
- Developers and artists seeking to generate more detailed and compositionally coherent images from simple text prompts.
- Users of ComfyUI looking for an integrated solution to enhance prompts and encode text within their Z-Image workflows.
- Anyone requiring a local, private solution for advanced image prompt engineering without relying on external APIs.