WISRD Benchmark Tests AI Image Editing Models with Visual IQ Tests
HirokatuKataoka · x · 2026-08-05
WISRD: Worksheet Image-Space Rule Discovery
The paper introduces WISRD, a novel end-to-end visual reasoning benchmark. It challenges image-editing models to solve problems entirely in image space—reading visual instructions, discovering rules, and writing answers directly onto the image without leaving the visual modality.
Core Challenges
The models must demonstrate multiple capabilities:
- Recognize problems and infer answers
- Bind answers to correct spatial destinations
- Control output count and suppress unnecessary edits
- Preserve the original input and format
Frontier Model Performance
- Nano Banana Pro achieved the highest score with a 48.7% pass rate on the no-reference subset.
- Qwen-Image-Edit (13.4%) and FLUX.2 Klein 4B (11%) showed mediocre performance.
- InstructPix2Pix scored 0.0%.
- In supplementary Sudoku and pattern reasoning diagnostics, Nano Banana Pro scored 70.0% and 22.9%, respectively.
Key Finding
Analysis reveals that current image-editing models can partially rely on rendered in-image instructions, even ignoring external text prompts.
More from Multimodal
- 14-Step Hand-Calculated Walkthrough of Sora's Diffusion Transformer Architecture — ProfTomYeh · 2026-08-05
- MiniMax Video Model Nails Complex Cinematic Prompts in Real-World Test — Devajyoti1231 · 2026-08-05
- Practical Workflow to Reduce Accents in Minor Languages for MiniMax H3 — martinerous · 2026-08-05
- Microsoft Launches MAI Image/Voice Models, Displacing OpenAI in Bing and PowerPoint — dl_weekly · 2026-08-05
- MiniMax H3 Video Model: High-Res Renders Nearly Identical to Low-Res, Revolutionizing Workflow — Far-Solid3188 · 2026-08-05
- Anime Short Made with Wan 2.2: Uncanny but Fun — JayoTree · 2026-08-05