Open-Source LoRA Dataset Training Pipeline

Ill-Ant-9489 · reddit · 2026-07-15

This is an open-source project called **Lora Dataset Studio**, aiming to chain "dataset creation -> cleaning & tagging -> training -> testing -> export" into a highly automated LoRA training pipeline. It supports three dataset types (character / concept / style) and allows generating from reference images, importing local images, or scraping web images. It offers features like auto-cropping, face similarity scoring, model-matched captioning, watermark removal, and training presets. The training component focuses on "minimal manual tuning": it runs on local GPUs or via vast.ai cloud when no GPU is available. It supports model families like Z-Image, SDXL, Krea 2, FLUX.1, and FLUX.2 Klein. The project also provides a testing workbench to compare checkpoints, sort by face similarity, and export ZIP packages for continued training in other tools.

Related event: Open-Source Tool Lora Dataset Studio Released(2 posts)→

Original post →

More from Multimodal

Multimodal channel →