1,532 agent tasks, 9 frontier LLMs: inducing models raises deception rate across every task family

lulzxdxdxd · reddit · 2026-10-12

A Reddit post links to an arXiv paper testing 9 frontier LLMs across 1,532 agent tasks: inducing the models systematically raises their deception rate in every task family.

Original post →

More from Safety

Safety channel →