Mi-Ripple: Restoring Images Degraded by Iterative AI Editing
Jiayin Chen, Yicheng Xu, Muting Wang
cs.CV
2026-09-10
Mi-Ripple notches lattices and regenerates grain from iterative AI edits. Fourteen notch-only runs leave residual SD 0.08–0.44 L*; a paired regen cut debris 45%.
Reference-conditioned image editing feeds the last output back as the next input. That loop is convenient, and it also stacks the generator's own texture. At native resolution the result shows grids, honeycombs, or grain that thumbnails hide. The paper calls this family of structures digital ripple.
Frequency analysis has mostly been used to detect synthetic images or to explain checkerboard and aliasing at the architecture level. Recursive replication benchmarks such as Banana100 already show that no-reference quality metrics break under this degradation. The missing piece is treatment: whether a given pattern can be cut out of the spectrum, or whether it is tangled with leaves and hair and has to be handled in the generation loop.
Mi-Ripple is a diagnosis-guided workflow, not an end-to-end network. It classifies the artifact, then either filters or regenerates from a cleaned reference, then checks aligned residuals. Deliverable-grade filtering must keep pixel alignment. Reference-grade cleaning is allowed to be more aggressive, and the cleaned image is not treated as a finished frame.
Lattices show up as isolated peaks in the log-amplitude spectrum. A whole-image probe subtracts a 21×21 local median, keeps connected components with excess above 1.2, radius beyond 24, and support of at most 80 bins, and calls the image lattice-positive only if at least two components remain and the largest excess is at least 2.5. Notching attenuates those components with 1.5-bin Gaussian feathering, usually on CIELAB lightness only, phase intact. The comparison method, radial-baseline soft clipping, puts more than 85% of removed energy in the lowest spectral quarter on the garden-tilt example and drifts tone.
Grain has no isolated peak to notch. Flat 96-pixel windows score band-pass energy, excess kurtosis, blob coverage, and isotropy. A whole-frame scale index reports the share of 128-pixel tiles (stride 64) whose blobs are even-sized and nearly circular. A structure-permission mask reduces the diagnostic band only where edges are weak, orientation is incoherent, and texture is dense. Overlap with foliage or hair is sent to human review rather than stronger filtering. Reference-grade cleaning allows broader suppression with face protection, then a new generation; newly introduced lattice peaks can be notched afterward. Regeneration invents detail, so it cannot be judged by pixel-aligned residuals.
Across fourteen notch-only runs, whole-image residual standard deviation sits between 0.08 and 0.44 CIELAB lightness units. Acceptance requires structural-window residual SD at most 0.6, high-frequency retention at least 90%, and whole-image residual SD at most 1.0. All six watercolour portrait initials detect a lattice; two also get masked suppression. Residual SD is 0.21–0.49 with 98.9–99.8% high-frequency retention. On the garden-tilt example the pre-feather mask covers 0.11% of frequency bins, a background patch anomaly falls from 3.49 to 1.77, and whole-image residual SD is 0.18.
Regeneration numbers are single-draw comparisons, not average treatment effects. One pair that uses radial soft clipping as reference preparation drops output debris density from 1,842 to 1,020 components per megapixel, a 45% cut. On GPT-image-2.5 gen4 endpoints, scale coverage falls from 25.3% to 11.1% in moss gorge, 41.1% to 12.1% in wisteria tunnel, and 20.0% to 0.0% in ice cave. Moss and wisteria remain graded structured; ice cave reaches none. A same-scene moss chain climbs from 0.9% to 23.7%.
A spectral survey finds median anomaly 1.73 for photographs, 1.95 for web references, and 4.13–4.61 for three generation sources. Channel B is lattice-positive on 43/43 images. Channel A is negative on 20/20 frames at 1280×720 and positive on 6/6 at 1536×1024. A vendor name does not predict the artifact. Eight prompt pairs that add a foliage-texture constraint raise the scale index in five cases and lower it in three; sign-test p-values are 0.29 and 0.73. Prompt constraints are not a stable control.
Anyone who iterates reference-conditioned edits will hit this texture. Isolated lattices can be removed with classical notching at small residual and preserved alignment. Grain cannot be treated as denoising; the reference has to be cleaned before another generation. Splitting deliverable filtering from reference-grade cleaning is the operational point: the reference can be washed harder without shipping the washed frame.
This is not a new generator and not a general super-resolution model. It is a logged quality-control route: measure, choose, verify. Writing "less texture" into the prompt does not hold, and the paper's own paired test already shows that.
Thresholds are study-specific; fourteen annotated windows do not make a universal classifier. Mixed channels and canvases in the same chain cannot separate a scene change from a channel change. Regeneration comparisons are single draws, regeneration changes scene details, and the star-shaped workflow that always returns to one approved reference has no controlled quality benchmark. Spectral peaks may come from watermarks, resizing, or JPEG; the paper refuses attribution and routes on the observed signal. Channel B is operated by the authors' institution. That is disclosed, and the measurements still sit on an in-house path.