LoRA Training with SAM 3D for Accurate Full-Body Proportions
tekprodfx16 · reddit · 2026-07-14
A developer forked the AI Toolkit project and integrated Meta's SAM 3D Body model, solving the issue of body type drift (height, proportions) during standard LoRA training for full-body generations.
Training Mechanism:
- Phase 1 (200 steps): Standard LoRA/LoKr training based solely on reference photos.
- Phase 2 (15-60 steps): After image generation, SAM 3D scans the body types of both generated and reference images, using a loss function to force the LoRA to learn the real person's 3D physical characteristics (rather than just guessing from text descriptions).
Performance & Compatibility:
- On an RTX 5090, training just the face takes 30 minutes; adding 3D body scanning takes 60 minutes.
- Currently supports the Krea 2 model; Ideogram 4 support is in development.
The author has open-sourced the project, significantly improving the realism and consistency of full-body character generations in AI art.
Related event: Enhancing LoRA Training with SAM 3D for Accurate Body Shapes(2 posts)→
More from Multimodal
- Pablo Stanley shares a full AI video workflow using ChatGPT, Gemini, Runway and CapCut — jdjohnson · 2026-07-21
- Meta AI text input now lets users interleave images with text — ezyang · 2026-07-21
- ShotPlan adds learnable planning tokens for cinematic multi-shot video generation — Tele-AI · 2026-07-21
- Same prompt, Seedance 2 and Grok are compared on cinematic transformation output — LudovicCreator · 2026-07-21
- CG Chefs Showcases Retro Anime Style AI Video Generation — nicolascraske · 2026-07-21
- Night-party video demo uses Seedance 2.0, timecode prompts and 4K upscaling — gen_ericai · 2026-07-21