192 local runs show even good models shouldn't authorize their own tool calls

Slight_Analysis_5414 · reddit · 2026-10-08

After 192 live local runs, the author argues models should propose tool calls but never authorize them: "models propose, systems enforce."

Code and benchmarks are open-sourced on GitHub.

Original post →

More from coding & agent

coding & agent channel →