yinggzhang/WeGenBench-Consistency-COT

VISIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 18, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The yinggzhang/WeGenBench-Consistency-COT is an 8 billion parameter text-to-image consistency evaluation model, fine-tuned from Qwen3-VL-8B-Instruct. It specializes in assessing how faithfully a generated image adheres to long, detailed prompts, providing a structured judgment with a 1-10 score and fine-grained deduction reasons. This model is primarily designed for automatic text-to-image evaluation within the WeGenBench pipeline, focusing on prompt-faithfulness rather than aesthetic quality.

Loading preview...

WeGenBench-Consistency-CoT: A Multimodal Judge for Text-to-Image Prompt Consistency

WeGenBench-Consistency-CoT is an 8 billion parameter model, fine-tuned from Qwen3-VL-8B-Instruct, specifically designed to evaluate the consistency between a generated image and its source prompt. Unlike general aesthetic judges, this model focuses exclusively on prompt-faithfulness, assessing whether the image accurately reflects the textual constraints.

Key Capabilities

  • CoT-style Deduction Grading: Provides a 1-10 score, total deduction, overall assessment, and detailed, interpretable deduction reasons for prompt-image mismatches.
  • Fine-grained Error Taxonomy: Identifies specific error categories such as entity, appearance, activity, counting, shape, material, text, and composition.
  • Long-Prompt Ready: Trained with a max_length=32768, making it suitable for complex prompts with numerous visual constraints.
  • Structured Output: Generates a parseable output format for automated analysis, detailing specific prompt requirements and actual image mismatches.

Good For

  • Automatic Text-to-Image Consistency Scoring: Ideal for objective evaluation of image generation models.
  • Cross-Model Comparison: Useful for pre-screening and comparing different text-to-image models based on their ability to follow instructions.
  • Badcase Mining and Data Filtering: Helps in identifying and categorizing images that fail to meet prompt specifications.
  • Regression Testing: Supports quality assurance by detecting consistency regressions in model updates.
  • Human Annotation Assistance: Provides detailed explanations that can aid human reviewers in understanding image generation errors.

This model is a critical component of the WeGenBench evaluation pipeline, offering a robust method for diagnostic analysis of text-to-image model performance.