Iris-3B scores 0.540 on OneIG, tying Qwen-Image, with no depth or 4× gain

Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning

Hanqiu Li Cai, Chema Garabito

SperidLabs

cs.CV, cs.AI

2026-10-07

Iris-3B, a 3B pixel-space text-to-image model, ties Qwen-Image at 0.540 on OneIG. On depth and 4× restoration, a pixel prior shows no clear gain over latent FLUX.2 Klein.

What problem this solves

Latent diffusion generates inside a VAE, which is why 1024-pixel training fits a realistic budget. The codec is trained on its own and throws information away, and the generator cannot put that information back. Depth and restoration then have to come out through a decoder fit to natural photos. On those tasks the codec is a structural limit, so a pixel-space prior ought to transfer better.

JiT and PixelDiT already approach latent models on the images they generate, and a separate line converts a finished latent model into pixel space at much lower cost. Those comparisons stop at generation or understanding. This report runs both routes. Iris-3B is a 3B text-to-image transformer pretrained from scratch on pixels. The second model is FLUX.2 Klein base 4B with the VAE removed. Both are fine-tuned for monocular depth and for 4× restoration. The transfer advantage does not appear.

Method

A 1.3B PixelDiT at 256 picked the recipe. Velocity prediction, with target equal to noise minus the clean image, trains stably and ties x-prediction on GenEval. x-prediction is worse on DPG and FID. JiT and the conversion literature favor x-prediction, but this x arm also used a noisier timestep schedule, so the target is not isolated. Among alignment losses, REPA to a frozen DINOv2 is the strongest overall. DINOv3 lowers GenEval and worsens FID. iREPA's FID lead has mostly disappeared by 250K steps, and GenEval and CLIP stay behind. Self-Flow is far behind on all four scores. The 3B run keeps velocity prediction and DINOv2, with the REPA weight at 0.5 only in the 256 stage.

Iris-3B matches flow on the straight path from image to noise, directly on RGB. Patches of 16×16 turn a 1024 image into a 64×64 token grid. The trunk is 8 dual-stream blocks plus 16 single-stream blocks at width 2560, with separate text and image weights in the dual-stream half, and GQA using 20 query heads and 5 key-value heads. One shared timestep projection per stream, plus a learned bias on each block, replaces per-block adaLN and drops 28% of those parameters at the same FLOPs. Four PiT blocks modulate each pixel from its patch token and, for attention, pack that patch into one 1280-d token. Prompts come from a frozen Qwen3-VL-4B-Instruct: 12 layers are pooled, then passed through a two-block refiner. 2D axial RoPE is not tied to resolution, so the curriculum reuses one set of weights. Hidden matrices use Muon and the remaining parameters use AdamW. The learning rate is 1e-4, then 5e-5 at 1024. The three stages see about 375.8M, 166.9M and 48.6M samples, plus 20.5M of supervised fine-tuning. Non-photos dominate at 1024, because fewer photographs clear the resolution cut.

Converting FLUX.2 Klein failed at the input layer. The trunk expects VAE embeddings. A random projection never produces that distribution, and the best linear map from a 16×16 RGB patch explains only 14% of the embedding variance. At 1K steps the samples are still prompt-independent patch noise. Two changes fix the probe when used together: a gain-matched ridge map into the parent's latent interface, and a trunk learning rate set to 0.12 times the head rate, because the FLUX trunk weights are much smaller than a normal Adam step assumes. On a batch of 8, the pair yields prompt-aligned photos within 1K steps. Either change alone cuts that 1K-step loss by more than half and still fails to make a coherent image. Production conversion uses batch 112. Downstream runs take the PiT head at 30K steps. GenEval, DPG and FID for the converted model were not measured.

Depth fine-tuning is one direct-regression recipe for all three parents. A single forward pass, with an empty prompt, zero noise and the final timestep, predicts relative log-depth. Pixel models read it with a convolution on RGB. The latent model first decodes with a frozen FLUX.2 VAE. The loss is masked L1 plus five times a gradient L1, for 10K steps at learning rate 3e-5 and batch 32, on Hypersim and Virtual KITTI 2 only. Regression had already beaten a flow objective on the converted parent. Restoration follows HYPIR's single step: upsample the degraded image, one pass, L2 plus five times LPIPS plus a 0.5 adversarial term. The FLUX pixel and latent arms share data, loss and schedule through 30K steps. The Iris restoration run does not. Its batch, schedule and wavelet color correction at inference all differ.

Results

Official evaluators score the released checkpoint at 1024, CFG 3, 100 steps and four samples per prompt. Prompts are not rewritten, and there is no best-of-N selection.

ModelGenEvalDPGLongTextOneIG
FLUX.1-dev 12B0.6683.80.6070.434
Qwen-Image 20B0.8788.30.9430.539
Z-Image 6B0.8488.10.9350.546
Iris-3B0.79886.520.8570.540

Iris-3B beats FLUX.1-dev on every column. OneIG is 0.540, level with Qwen-Image at 0.539 and just under Z-Image at 0.546. GenEval 0.798 and LongText 0.857 remain behind Qwen-Image at 0.87 and 0.943. i1 reports DPG 86.7, next to Iris at 86.52. i1 and Z-Image both list GenEval 0.84, and both rewrite prompts. Iris is strongest on single objects, two objects and color, and weakest on position and counting.

Depth is zero-shot under Marigold V2's scorer, which may fit a scale and shift per image. SEE is the edge-sensitive error on the Hypersim test split.

ModelMean AbsRel↓Mean δ1↑Hypersim SEE↓
Klein latent0.0720.9470.528
Klein pixel0.0810.9340.526
Iris-3B0.0710.9460.496

Iris and latent Klein tie on the mean: AbsRel 0.071 versus 0.072, and δ1 0.946 versus 0.947. Iris is better on DIODE (0.082 vs 0.107) and KITTI (0.083 vs 0.086), and worse on NYUv2 (0.057 vs 0.050). Converted pixel Klein averages 0.081 and loses on every benchmark except KITTI. These are single short runs. The per-set gaps cancel, and the only consistent deficit is the briefly converted model.

The same depth maps carry a patch grid. A ratio of border curvature to curvature elsewhere lands near 1 when no grid is present. Latent Klein scores 0.97 to 0.99. Pixel Klein scores 1.31 to 1.44. Iris reaches 2.14 on Hypersim and 2.25 on DIODE. Non-overlapping patches are decoded mostly from their own trunk token, so neighbors can disagree at the cut. A VAE decoder overlaps those tokens and blends the seam. The wiggle is smaller than the depth error, and it lands on the fine structure pixel space was supposed to preserve.

On 4× DIV2K with the dataset's native Real-ESRGAN degradations, latent Klein scores 20.99 PSNR, 0.557 SSIM, 0.275 LPIPS, 3.26 NIQE and 67.11 MUSIQ. Pixel Klein scores 20.87, 0.504, 0.274 and 67.03 MUSIQ. SSIM and NIQE are the clear losses. LPIPS is flat at 0.274 versus 0.275. Iris scores 20.68, 0.481 and 0.292, with NIQE 3.94 between the two Klein runs, DeQA 3.87 slightly below both, and MUSIQ 68.01, the highest value in the table. SeeSR still leads fidelity, at 22.25 PSNR and 0.580 SSIM. Real-ESRGAN's MUSIQ is 60.37. The pixel restorations sit with the stronger open systems on no-reference scores, and they do not beat the latent twin.

Why it matters

A PiT head and a from-scratch curriculum can train a 3B pixel model to 1024 without a VAE. On the official OneIG evaluator, that 3B model ties 20B Qwen-Image. Weights and training code are released. The 28% parameter cut from shared-bias adaLN, and RoPE with no resolution-specific parameters, can be copied as they stand.

The dense-prediction claim does not hold. Under the matched recipes, a pixel-space generative prior brings no visible gain. Dropping the lossy codec does not, on these runs, produce better depth or better 4× restoration. The report flags editing as the next test: unchanged regions have to be reproduced pixel for pixel, which is where a lossy latent should cost the most. That experiment is not in the paper.

Limitations

Iris has no latent twin, so model size, data and compute are mixed into the depth comparison with FLUX.2 Klein. Both matched pairs share a parent that was converted for only 30K steps, and GenEval, DPG and FID for that parent were never published. A weaker prior is a plausible reading, not a measurement. Every downstream arm is one short run, with no significance test. Iris restoration also changes the batch, the schedule and the postprocessing. The 1024 pretraining mix is mostly non-photographic, and nothing in the report checks whether that dilutes the real-image detail prior.

The x-versus-v ablation moves the timestep distribution together with the target. Ridge initialization and the scaled trunk learning rate were probed at batch 8 for 1K steps, not at the batch of 128 used by published conversions. The PiT head never lets pixels on opposite sides of a patch border meet at full resolution, so the grid is structural. It works against the spatial detail that motivated pixel space. Long text, position and counting remain the open gaps on the generation side.

Terms

Source

What people are saying

Related papers

All paper explainers