2026-09-03
Anamorph1x LoRA restores oval bokeh that diffusion priors sand off, at strength 0.8 on a 1024×504 canvas in a dual-sampler ComfyUI plus SEEDVR2 pipeline.
Generative image systems treat quality as total visibility: resolution, sharpness, denoise, spatial coherence. Streaming did the same to cinema, pushing living-room brightness and HDR until grain looked like a compression tax. The resulting video look is information-rich and film-poor.
Anamorphic capture is the counterexample. Oval bokeh, edge stretch, compressed depth, and center-weighted sharpness are how those lenses organize space. Diffusion models trained on internet-scale clean photos treat those signatures as noise. Prompting "cinematic anamorphic" usually buys a wide canvas and a little blur; the skeleton stays evenly detailed. Without changing the prior, prompts are decoration.
Most of the paper is optical and historical argument, not a new trainer. Cases run from Ridley Scott's Legend and Corman's Poe cycle through Tarkovsky's Solaris, Xavier Dolan's 4:3 Laurence Anyways, and Dogma 95. Format and lens have to be locked before capture. Desaturation or grain overlays in post cannot retrofit an optical grammar. The 4:3 frame is used as a control case: shrinking the canvas changes the compositional logic, the same way anamorphic glass does. It is a language, not a LUT.
On the generative side it lists three operating rules:
The concrete artifact is anamorph1x, a LoRA the author trained on a custom Z-image diffusion stack. LoRA is a small low-rank patch on a frozen base model. The token is not the whole trick; the weights push oval vertical bokeh, mild edge distortion, and horizontal composition. Version 2 used stills with genuine anamorphic traits. A later version is planned on frames the author shot with real anamorphic glass, and on FLUX and Wan Diffusion 2.2. Strength was swept from about 0.6 to 1.0; most outputs that felt anamorphic sat near 0.8. Canvas was 1024×504, about 2.035:1. The ComfyUI graph IAMCCSZ-IMAGEanamorph1x runs a dual k-sampler: compose small, 2× upscale, refine, then a SEEDVR2 pass to keep the optical character.
No numbers that would survive a methods review. No FID, no CLIP score, no user study, no prompt-only ablation. What exists is a recipe and an author's eye: strength near 0.8 on a wide canvas. The paper itself calls anamorph1x a humble first step.
The reproducible bits are the file name anamorph1x.safetensors, the ComfyUI workflow name, the strength band, 1024×504, and the SEEDVR2 finish.
For people who generate moving-image frames, the useful sentence is this: the model's default good photo is not a cinematographer's default good frame. If you want oval bokeh, train an optics LoRA instead of stacking the word cinematic. The library idea is practical: one adapter for austere 4:3, one for 16mm grain, one for Technicolor, chosen the way a set chooses a lens package.
This is not a new training method. LoRA, reference conditioning, and multi-pass sampling are off-the-shelf. The selection is the point. Optical imperfection is a generation target, not a defect. Worth trying. Not evidence that cinematic AI has arrived.
The quantitative hole is the real problem. "Mostly works at 0.8" has no sample size and no failure census. Dataset size, steps, rank, and learning rate are missing. The ethically constrained author-driven pipeline is the author picking frames and wiring a graph; replication is weak.
The film-history sections are criticism, not measurement. Scorsese on Marvel as theme parks and Nolan on celluloid do not, by themselves, prove streaming killed film texture. Section 7 admits unmanaged anamorphic looks ugly, yet the generative half never shows those failures. Z-image, FLUX, and Wan 2.2 sit in one pipeline paragraph, so it is unclear which stage actually carries the anamorphic look.