1,680 A/B trials show 50-year-old Unix formalisms crush English prompt rules, 98.3% vs 0.8%

RandalSchwartz · reddit · 2026-10-11

The author ran 1,680 paired A/B trials (41M tokens, four model tiers, each loaded with 16k tokens of compiler noise) pitting 500-token English prompt rules against 20-token classic CS formalisms, and found that custom prompt DSLs (pre-training frequency ≈ 0) force the model to simulate a fragile interpreter, while Unix/RFC formalisms (frequency > 1M) snap it into strict compliance.

Key results:

Takeaway: reuse Makefiles, init-d runlevels, and fork()/wait() subagent isolation instead of inventing English checklists the model has never seen.

Original post →

More from coding & agent

coding & agent channel →