Qwen3-8B fine-tuned with LoRA+GRPO mimics a famous ML blogger, fooling every AI detector

OtherRaisin3426 · reddit · 2026-08-28

The author fine-tuned open-source Qwen3-8B-Base in two stages — LoRA SFT then GRPO — on 390k words of a famous ML author's public posts, and built a writing Turing test: of 10 passages, half are real and half model-generated.

Key result: every commercial AI detector they tried clears the generated passages (Pangram flagged 0/10). The GRPO reward combines a style discriminator with a pairwise judge against real passages, and the discriminator is retrained on fresh rollouts so the model can't hack a frozen classifier.

Playable at turing-writing-test.vercel.app.

Original post →

More from Models

Models channel →