How to fine-tune an LLM for personal RP/writing style? Guide on dataset and SFT+DPO workflow
ba2sYd · reddit · 2026-08-20
A user seeks advice on fine-tuning an LLM (like TheDrummer's models) for personal role-playing (RP) and writing style. The proposed plan involves training on personal RP logs and fiction, potentially adding a DPO step with a preference dataset. Key questions include: the required number of samples, optimal dataset structure, whether SFT+DPO is the best approach, recommended hyperparameters, and caveats when training on personal writing.
More from coding & agent
- Practical Prompt: Agent Workflow for Auto-fixing PR Review Comments — iannuttall · 2026-08-20
- Don't just build a chat box: Real Agent value lies in workflows — sujingshen · 2026-08-20
- Paste an OpenAPI spec, get a working MCP server — free platform seeks testers — Own_Initial_670 · 2026-08-20
- User Suggests Integrating Grok Build Features into Cursor CLI — intellectronica · 2026-08-20
- Agent tracks Ontario lakes, reverse engineers realtor sites for leads — yacineMTB · 2026-08-20
- Grok Bot Demo: Automating Business, Software, and Marketing via X and GitHub — elonmusk · 2026-08-20