DEPICT: Training-Free Alignment Metric Boosts Negation Accuracy from 19% to 88%

swordhealth · hf · 2026-10-06

Sword Health proposes DEPICT, a training-free image-text alignment metric addressing gaps in fine-tuned (backbone-bound), holistic (misses details), and decomposed (fixed-YES assumption) evaluators.

Method: replace fixed reference answers with expected agreement between image-based and caption-only answers, weighting questions by how decisively the caption determines them, then merge with a holistic score to recover lost context.

Results: the agreement rule raises negation accuracy from 19% to 88%; across 5 benchmarks and 11 backbones from 3 model families, DEPICT beats all training-free metrics and exceeds fine-tuned evaluators on 2 of 3 human-correlation benchmarks.

Original post →

More from Multimodal

Multimodal channel →