RISEBench++: 65 reasoning-based visual editing tasks; best model GPT-Image-2.5 hits only 56.6%
VisionXLab · hf · 2026-10-09
- RISEBench++ extends the first reasoning-informed visual editing (RISE) benchmark: six reasoning dimensions (temporal, causal, spatial, logical, counterfactual + hybrid), 12 subcategories, 65 fine-grained task types, multi-image conditioning, and 1000 human-annotated bilingual test cases.
- Evaluates 58 approaches (34 open-source, 19 closed-source, 5 agentic) on instruction reasoning, appearance consistency, and visual plausibility with human and LMM-as-a-judge scoring.
- Findings: reasoning-based editing remains far from solved—the strongest, GPT-Image-2.5 Sunburst, reaches only 56.6% accuracy. Also introduces RISE-Agent, a training-free agentic framework (planning + tools + verifier-guided refinement) that beats most strong baselines.
More from Multimodal
- Autoregressive Retriever (ARR) Refines Queries with Retrieved Item Feedback via SFT and RL — _reachsumit · 2026-10-09
- Sony's Syn-Omni: Shared + Expert LoRA Paths Beat Omnimodal Embedding Baselines Across 81 Tasks — _reachsumit · 2026-10-09
- LEGO: lifting-free exocentric-to-egocentric video generation beats depth-lifting SOTA pipelines — 25frms · 2026-10-09
- VibeEdit Replaces Text Prompts with Canvas Marks, Scoring 79.9 on Edit Benchmark — Sydney-Uni · 2026-10-09
- Microsoft's Compo Shifts Poster Generation from Prompting to Spatial Composing — microsoft · 2026-10-09
- Open-AI-Design-Agent: a free MIT-licensed open-source alternative to Lovart, Runway and Luma design agents — matchaman11 · 2026-10-09