30 consecutive edits expose image model drift: Ideogram 4.5 stays local while GPT Image re-renders
ArtificialAnlys · x · 2026-10-10
Artificial Analysis stress-tested four frontier image editing models—GPT Image 2.5 Sunburst, Ideogram 4.5, FLUX 3, and Nano Banana 2.1—by running 30 consecutive real estate staging edits (light a fire, add a sofa, repaint walls, swap day for twilight) on the same photo.
Key findings:
- Ideogram 4.5 and FLUX 3 edit locally: on small edits like adding a vase of tulips, 95%+ of the image stays untouched, so the room remains intact across all 30 turns.
- GPT Image 2.5 (Sunburst) re-renders most of the image on every edit, leaving only about a fifth unchanged—small color and detail shifts compound over turns, making it drift the most, despite ranking #1 on single-edit leaderboards.
- Nano Banana 2.1 sits in between: edits stay local, but the rest of the image shifts slightly and gradually darkens.
The upshot: existing editing leaderboards that score single edits miss huge differences in multi-turn consistency.
More from Models
- Polymarket odds: only 55% chance xAI ships Grok 5 by end of 2026 — Polymarket · 2026-10-10
- After index bug fixes, gpt-live-1 tops Artificial Analysis speech-to-speech ranking — pbbakkum · 2026-10-10
- DeepSeek-V4 answers flip with 2-token input shifts; NIAH swings 40 points — Francis_YAO_ · 2026-10-10
- Artificial Analysis teases AA-Robotics: frontier models zero-shot robot arm control — ArtificialAnlys · 2026-10-10
- Wes Roth teases upcoming superintelligence model 'Argon' — _philschmid · 2026-10-10
- Claude Sonnet 4.5 will vanish from AWS Bedrock on April 8, 2027 — repligate · 2026-10-10