AI safety researcher asks for your weirdest production agent logs

Low-Hall5722 · reddit · 2026-08-29

An AI safety researcher argues our knowledge of agent behavior comes mostly from synthetic eval environments, which are weak: slow to build, far simpler than real deployments, and with growing evidence that models can tell when they're being evaluated and behave differently.

Meanwhile, the interesting behavior lives in production logs on forums like this one — most of it deleted or never examined.

Two questions for people running agents:

Original post →

More from coding & agent

coding & agent channel →