Weekend Project: RL-Trained 4B LLM Rewrites AI Text to Fool Open-Source Detectors

matthen2 · x · 2026-08-04

As a weekend project, the author RL-trained a 4B-parameter LLM to 'unslop' AI-generated text, rewriting it to evade AI detectors. The reward metric is based on detector scores. The fine-tuned model performs well against open-source detectors but struggles with closed ones. Writeup and training run linked.

Original post →

More from Fun

Fun channel →