Reasoning-Informed Visual Editing
Xue Yang, Peiyuan Zhang, Yilun Zhu, Qihao Yang, Mingxin Liu, Xiangyu Zhao, Ziqian Fan, Zhaokai Wang, Yan Li, Yifan Yang, Xu Yang, Xiaosong Jia, Yue Zhou, Zhihang Zhong, Junchi Yan
cs.CV
2026-10-09
RISEBench++ tests 58 editors on 1,000 bilingual reasoning edits. GPT-Image-2.5 Sunburst leads at 56.6%; logical accuracy falls to 27%, and the best open model hits 23%.
Editors already recolor, swap objects, and relight a scene. They drop once the prompt never states the target pixels and the model has to infer them: a reaction ten minutes later, or the camera's position marked on a map. Diffusion editors spend their capacity on appearance. Inferring the target from knowledge, layout, or a hidden rule was never trained as its own skill.
Benchmarks of explicit commands miss that skill. The conference RISEBench, an oral at NeurIPS 2025 Datasets and Benchmarks (top 0.35%), had 360 items, four reasoning types, and fewer than ten models.
RISEBench++ is a three-level set: six dimensions, 12 subcategories, 65 task types, and 1,000 human-annotated items in English and Chinese. Temporal, causal, spatial, and logical reasoning have 200 items each. Counterfactual and hybrid reasoning have 100 each. A hybrid item runs three to five turns, and each turn edits the previous output.
Temporal and causal items are mostly knowledge-driven. With the right facts, the target state is a short inference. Spatial and logical items are perception-driven: the layout or the rule has to be read off the picture. Counterfactual items keep commonsense and then break it on purpose.
Three scores run from 1 to 5 and are reported on a 100-point scale. Instruction reasoning uses a reference caption for simple scenes and a reference image for logical and spatial items that text cannot pin down. Appearance consistency checks regions the instruction did not ask to change. Natural photos get a graded score. Logical items, often sparse diagrams, are binary: 5 if untouched content survives, 1 if it does not. Visual plausibility looks only at edited regions. Counterfactual items are not penalized for breaking physics. An item is solved only when all three scores are 5. Accuracy is the share of solved items. Gemini-3-Flash judges the first two axes. Gemini-3.1-Flash-Lite judges plausibility.
RISE-Agent adds no training and splits inference from painting. A GPT-5.1 planner reasons directly, searches the web, or calls seven code solvers covering arithmetic, dates, shortest paths, Sudoku, coordinate geometry, and matrix transforms. The executor is open FLUX.2 Klein 9B plus 12 deterministic operations: lines, polygons, text, recoloring, and moves. Pixels outside the operation stay put. A Gemini-3-Flash verifier picks pass, local repair, or replan, with at most one replan and two repairs.
On English, GPT-Image-2.5 Sunburst scores 56.6%, Nano Banana Pro 54.6%, and Nano Banana 2 50.8%. GPT-Image-2, GPT-Image-2.5 Flare, Luma Uni-1, and Qwen-Image-3.0 sit near 45%. The best open model, HunyuanImage-3.0-Instruct, reaches 23.0%. Qwen-Image-Edit-2511 reaches 17.1%. On single-image items, RISE-Agent scores 48.3%, MindBrush 21.8%, InterleaveThinker 20.6%, IntentEdit 13.2%, and ImAgent 3.5%. RISE-Agent's full-set score is 48.0%. The other agents skip multi-image inputs, so they have no full-set number.
Sunburst scores 67.5% temporal, 69.0% causal, 64.0% spatial, and 72.0% counterfactual, then 27.0% logical and 39.0% hybrid. Nano Banana Pro hits 71.0% causal, 36.0% logical, and 38.0% hybrid. Implicit pattern induction tops out at 31.2% among closed models, on Qwen-Image-3.0. Across seven models, pattern prediction never exceeds 25.6%. RISE-Agent reaches 47.2% on explicit rule deduction, above Sunburst at 32.5%, and 39.0% logical overall, above Sunburst at 27.0%. Spatial is 37.5% and hybrid is 28.0%.
Chinese order stays close for the leaders. Sunburst scores 55.5% and Nano Banana 2 scores 51.0%. GPT-Image-1 falls from 31.8% to 27.2%. RISE-Agent falls from 48.0% to 40.3%, and its logical score falls from 39.0% to 25.0%.
Newer models already bunch up on appearance consistency and visual plausibility. Instruction reasoning is where they separate.
The ablation uses the same accuracy. FLUX.2 Klein 9B alone scores 12.9%, with logical at 0.5%. The planner alone lifts the total to 40.5%, temporal from 16.0% to 58.0%, and causal from 16.5% to 50.5%, while logical is still 14.0%. Tools then lift logical to 32.0%, but temporal falls to 50.5% and spatial falls from 48.0% to 37.5%. The verifier brings temporal back to 58.5%, logical to 39.0%, and hybrid from 12.0% to 28.0%. Spatial stays at 37.5%. The total is 48.0%.
Judge agreement used 100 items, two models, 200 images, and four annotators. Gemini-3-Flash matches humans on instruction reasoning at 0.48 MAE and 0.84 Spearman. On plausibility, Gemini-3.1-Flash-Lite has 0.27 MAE and only 0.41 Spearman.
| Method | Metric | Result |
| GPT-Image-2.5 Sunburst | English accuracy | 56.6%, logical 27.0% |
| Nano Banana Pro | English accuracy | 54.6% |
| RISE-Agent | English accuracy | 48.0%, explicit rules 47.2% |
| HunyuanImage-3.0-Instruct | Best open, English | 23.0% |
| FLUX.2 Klein 9B plus full agent | Ablation | 12.9% to 48.0% |
Everyday causal edits are no longer the hard part. GPT-Image-1.5 scores 80.0% on causal common sense, Sunburst 78.8%, and Nano Banana 2 scores 77.6% on temporal common sense. Finding a rule the prompt never states still drops closed models to around 30%. The data to collect next is implicit rules and multi-turn edits. More recoloring pairs will not move this table.
On shortest paths, Sudoku, and grid flips, a generator that misses one symbol or one edge fails the whole item. RISE-Agent computes the target with a solver and paints it with deterministic ops, which is why its logical score beats pure generators. The 48.0% belongs to a GPT-5.1 planner on an open editor. Klein 9B alone is at 12.9%.
Against the conference version, the set grows from 360 to 1,000 items and the pool to 58 models, with multi-image, multi-turn, and Chinese added. Reasoning is still the bottleneck.
A solve requires a perfect 5 on every axis. Models that already hold appearance and plausibility still fail the item when the reasoning is slightly off. 56.6% is a perfect-execution rate. A one-cell miss and a wholly wrong plan look the same, and the near-miss rate is not reported.
Human agreement covers two models, not 58. Plausibility correlation is 0.41 Spearman, weak for ranking. The instruction-reasoning judge and RISE-Agent's verifier are both Gemini-3-Flash. The agent has already repaired outputs toward that judge, so 48.0% may sit high.
Tools are not free. After they are added, temporal accuracy falls from 58.0% to 50.5% and spatial accuracy from 48.0% to 37.5%. The verifier never recovers spatial. The table does not show which calls were mistakes.
Language is uneven. Sunburst moves from 56.6% to 55.5%. RISE-Agent moves from 48.0% to 40.3%. The planner may also search the web, and rerun stability is unreported.
High counterfactual scores follow the rubric. Sunburst's 72.0% counterfactual exceeds its own 67.5% temporal. Physics violations are not penalized, so square tires are closer to a stylistic rewrite. A logical item fails if one cell is wrong. The six headline percentages are not on one difficulty scale.
Some open models have no multi-image score, and comparing them with Sunburst's 56.6% full-set number overstates the gap. Head to head on single images it is still 56.9% versus 22.5% for HunyuanImage-3.0-Instruct.