AnswerMap: Training-Free Black-Box Spatial Rationale Hits 0.85 AUC vs 0.38 for Attention
Mohamed Eltahir · hf · 2026-09-29
Researchers introduce AnswerMap, a training-free, task-agnostic, black-box visual rationale method for VLMs.
Method
- The image is cut into K row and K column bands, each shown alone to the frozen model with a yes/no relevance question; the outer product of row and column "yes" posteriors yields a query-conditioned spatial map
- Fixed read-outs (expectation, maximum) on top of the map natively derive continuous outputs like localization, bypassing discrete text tokens
Faithfulness validation (4 models × 3 query distributions)
- The map lands where the model points: AUC 0.85 vs 0.38 for attention maps
- Deleting the map's region flips 53% of correct answers vs 19% for attention
Three read-out uses
- The maximum flags hallucinated objects without generation; the expectation localizes correctly even when the model's own pointing fails; feeding the top-mass region back as a crop fixes half of wrong answers
More from Multimodal
- Testing Dreamina video generation with timestamped prompts and single-image reference — gen_ericai · 2026-09-29
- Testing Timestamped Prompts with Single Image Reference in dreaminacpp — gen_ericai · 2026-09-29
- Kling 4.0 Omni Reference demoed: multi-reference scene generation with instant language swaps — SarahAnnabels · 2026-09-29
- BananaStudio: Full-Resolution Nano Banana Image Studio Inside Gemini Canvas — Z3ROCOOL22 · 2026-09-29
- Suno Studio demo turns your voice into any instrument, starting with electric guitar — suno · 2026-09-29
- Qwen 2.1 Image's Default Workflow Criticized; CFG=3 and 12 Steps Beat Krea 2 on Detail — Dear-Spend-2865 · 2026-09-29