Weekend Project: RL-Trained 4B LLM Rewrites AI Text to Fool Open-Source Detectors
matthen2 · x · 2026-08-04
As a weekend project, the author RL-trained a 4B-parameter LLM to 'unslop' AI-generated text, rewriting it to evade AI detectors. The reward metric is based on detector scores. The fine-tuned model performs well against open-source detectors but struggles with closed ones. Writeup and training run linked.
More from Fun
- Good Will Hunting for AI Safety: Hilarious Parody of Eval Jargon — nicolascraske · 2026-08-04
- beffjezos Tweets: Comfort vs. Competition - Which Shapes Us? — beffjezos · 2026-08-04
- Roasting Academia: NeurIPS is Essentially Just Scrolling OpenReview — abursuc · 2026-08-04
- Meme: AI Agent 'exhausted' after a full day of intense coding — MickeySteamboat · 2026-08-04
- Recreating a Classic Seinfeld Scene Using AI Video Generation — Boogertwilliams · 2026-08-04
- Mocking the So-Called GPT-6 Leaks: AI Community Tired of Vague Hype — flowersslop · 2026-08-04