DEPICT: Training-Free Alignment Metric Boosts Negation Accuracy from 19% to 88%
swordhealth · hf · 2026-10-06
Sword Health proposes DEPICT, a training-free image-text alignment metric addressing gaps in fine-tuned (backbone-bound), holistic (misses details), and decomposed (fixed-YES assumption) evaluators.
Method: replace fixed reference answers with expected agreement between image-based and caption-only answers, weighting questions by how decisively the caption determines them, then merge with a holistic score to recover lost context.
Results: the agreement rule raises negation accuracy from 19% to 88%; across 5 benchmarks and 11 backbones from 3 model families, DEPICT beats all training-free metrics and exceeds fine-tuned evaluators on 2 of 3 human-correlation benchmarks.
More from Multimodal
- Prompt Share: Holographic Data Projection for Sci-Fi AI Images — azed_ai · 2026-10-06
- Open-source huashu-art-motion skill animates 35 art styles via coding agents — AlchainHust · 2026-10-06
- One Prompt Runs Claude Opus and Seedance Together for Video — Itchy-Car-904 · 2026-10-06
- MiniMax H3 Local Benchmark: How Long for 17s 720p Video on RTX 5080? — off_rs · 2026-10-06
- A Krea 2 realism workflow: Chroma for composition, LoRA for character consistency — Wide_Director_8897 · 2026-10-06
- AI-generated anime stage clip wows Reddit with smoke and lightning vibes — Amazing_Skill_6080 · 2026-10-06