Faces break in SD1.5 LoRA trained on 500 uncurated images — what went wrong?

Available_Fondant_11 · reddit · 2026-09-08

A student building a pipeline for illustrated health-education stories trained an SD1.5 LoRA (rank 8, lr 1e-4, 1500 steps, 512px) on 500 openly-licensed images pulled via the Openverse API, auto-captioned with BLIP plus a trigger word.

The problem: backgrounds and scenes render fine, but faces are warped and uncanny. Layering on a comic checkpoint reproduced it, and img2img without the LoRA preserves faces — pointing at the LoRA training itself.

Working theory: uncurated data full of small/distant faces, profiles, group shots, and low-res crops. The author asks for sanity checks on other common causes (rank/step ratio, overfitting on noisy data, BLIP caption quality) and best practices for face-heavy LoRA datasets.

Original post →

More from Multimodal

Multimodal channel →