UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation
Abstract
Multi-modal image generation, particularly subject-driven customization, has garnered growing attention in recent years. Despite the rapid advancement of generative models, their evaluation remains largely lagging. Existing methods, whether embedding-based or Multi-modal Large Language Model (MLLM)-based, evaluate alignment with each modal condition in isolation, which contradicts the simultaneous condition alignment objective of multi-modal image generation, leading to poor consistency with human judgments. To address this challenge, we propose UFO, the first unified framework for omni-condition alignment simultaneous evaluation. Specifically, UFO introduces a novel Atomized Chain-of-Evaluation paradigm, i.e., it first decomposes omni-condition alignment into a sequential chain of fine-grained, disentangled Atomic Evaluation Units (AEUs), categorizes them into distinct modality-relevance classes, and then employs general or dedicated functional calls for accurate verification of different AEU types. Experimental results demonstrate that UFO achieves the highest correlation with human evaluation preferences, delivering an average improvement of 15.25%. Furthermore, we present UFO-Bench, a dedicated benchmark designed to holistically evaluate the performance of existing customization models under the diverse mutual interactions of textual and visual conditions.
Community
Multi‑modal image generation, particularly subject‑driven customization, has garnered growing
attention in recent years. Despite the rapid advancement of generative models, their evaluation
remains largely lagging. Existing methods, whether embedding‑based or Multi‑modal Large
Language Model (MLLM)‑based, evaluate alignment with each modal condition in isolation,
which contradicts the simultaneous condition alignment objective of multi‑modal image
generation, leading to poor consistency with human judgments. To address this challenge,
we propose UFO, the first UniFied framework for Omni‑condition alignment simultaneous
evaluation. Specifically, UFO introduces a novel Atomized Chain‑of‑Evaluation paradigm, i.e., it
first decomposes omni‑condition alignment into a sequential chain of fine‑grained, disentangled
Atomic Evaluation Units (AEUs), categorizes them into distinct modality‑relevance classes, and
then employs general or dedicated functional calls for accurate verification of different AEU
types. Experimental results demonstrate that UFO achieves the highest correlation with human
evaluation preferences, delivering an average improvement of 15.25%. Furthermore, we present
UFO‑Bench, a dedicated benchmark designed to holistically evaluate the performance of existing
customization models under the diverse mutual interactions of textual and visual conditions.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- VicEdit: Learning to Edit Videos from Visual In-Context Examples (2026)
- TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation (2026)
- MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing (2026)
- Evaluation-Verification Reward for Consistent Multi-Reference Image Editing (2026)
- GenPuzzle: Benchmarking Visual Reasoning in Image Generation Models (2026)
- A Model-Internal Protocol for Assessing Multimodal Models as Integrated Systems (2026)
- CPI-Bench: A Comprehensive, Practical and Intelligent Benchmark for Real-World Image Editing (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper