Faces break in SD1.5 LoRA trained on 500 uncurated images — what went wrong?
Available_Fondant_11 · reddit · 2026-09-08
A student building a pipeline for illustrated health-education stories trained an SD1.5 LoRA (rank 8, lr 1e-4, 1500 steps, 512px) on 500 openly-licensed images pulled via the Openverse API, auto-captioned with BLIP plus a trigger word.
The problem: backgrounds and scenes render fine, but faces are warped and uncanny. Layering on a comic checkpoint reproduced it, and img2img without the LoRA preserves faces — pointing at the LoRA training itself.
Working theory: uncurated data full of small/distant faces, profiles, group shots, and low-res crops. The author asks for sanity checks on other common causes (rank/step ratio, overfitting on noisy data, BLIP caption quality) and best practices for face-heavy LoRA datasets.
More from Multimodal
- Seedance 2.5 workflow demo: render from 3D scene, then edit conversationally — ZabihullahAtal · 2026-09-08
- One creator batch-generated 1,000 music videos from a single song via ComfyUI — Lucky_Feedback9915 · 2026-09-08
- Stable ComfyUI on RX 9070 XT: 12-16s per image after warmup, full config shared — OrionNebulae · 2026-09-08
- Kling AI heads to TIFF Market 2026 with forum on cinematic AI video — aziz4ai · 2026-09-08
- Fallout-trained H3 world model demo: WASD control plus LLM-reactive combat — -Ellary- · 2026-09-08
- AI video's paradox: better generation makes production flaws more obvious — ZabihullahAtal · 2026-09-08