geocine/minimax-video-prompt-enhancer-350m
The geocine/minimax-video-prompt-enhancer-350m is a 350 million parameter model fine-tuned from LiquidAI/LFM2.5-350M. It functions as a prompt rewriter, specifically designed to transform rough video ideas into structured MiniMax H3 video prompts, including details for shots, camera, soundscape, and score. This model excels at generating precise, formatted text prompts for various video generation modes like T2VA, I2VA, FL2VA, L2VA, and full-reference tasks.
Loading preview...
MiniMax Video Prompt Enhancer (LFM2.5-350M)
This model, fine-tuned from LiquidAI/LFM2.5-350M, specializes in enhancing rough video ideas into structured MiniMax H3 video prompts. It acts as a prompt rewriter, generating detailed specifications for shots, camera movements, soundscapes, and musical scores, adhering to the exact field layout expected by MiniMax H3.
Key Capabilities
- Structured Prompt Generation: Converts free-form video concepts into highly structured prompts for various MiniMax H3 modes.
- Support for Multiple Video Modes: Handles
T2VA(text-only),I2VA(first-frame),FL2VA(first + last frame), andL2VA(last-frame) tasks. - Full-Reference Rewriting: Supports advanced full-reference tasks such as
reference_generation,keyframe_completion,video_editing, andvideo_continuation, including audio referencing. - Specific Output Formatting: Generates output in either a three-field layout (integrated_multimodal_description, overall_soundscape, non_diegetic_music) for base tasks or a six-section layout (subject_definitions, summary, retention_analysis, detailed_description, overall_soundscape, non_diegetic_music) for full-reference tasks.
- ChatML Compatibility: Expects input in ChatML format with a system prompt and a user message containing task, duration, assets, and the user's rough idea.
Good for
- Users needing to generate highly specific and formatted video prompts for the MiniMax H3 platform.
- Developers integrating video prompt enhancement into applications that require structured text outputs for video generation.
- Automating the creation of detailed video production briefs from simple textual descriptions.
Limitations
- Not a Chat Model: Optimized specifically for prompt rewriting, not general-purpose conversation.
- MiniMax H3 Specific: Output format is tailored for MiniMax H3, limiting direct applicability to other video generation systems.
- Small Model Size: May require light post-processing for very long or complex multi-shot briefs.
- Text-Only Output: The model generates text prompts; it does not generate video content itself.