How to fine-tune an LLM for personal RP/writing style? Guide on dataset and SFT+DPO workflow

ba2sYd · reddit · 2026-08-20

A user seeks advice on fine-tuning an LLM (like TheDrummer's models) for personal role-playing (RP) and writing style. The proposed plan involves training on personal RP logs and fiction, potentially adding a DPO step with a preference dataset. Key questions include: the required number of samples, optimal dataset structure, whether SFT+DPO is the best approach, recommended hyperparameters, and caveats when training on personal writing.

Original post →

More from coding & agent

coding & agent channel →