Multi-LLM social worlds are 'large': models may learn heuristics instead of grader-modeling

1a3orn · x · 2026-09-11

Continuing his discussion with jdpressman, 1a3orn argues multi-LLM training is probably complex enough that rather than fully modeling the grader, models might instead learn something like desired heuristics, since social worlds are 'large' worlds.

The author admits he isn't really sure, but the point echoes his earlier argument that environment design determines whether models solve problems or obsess over the grader.

Related event: Researcher Questions Whether Bad RL Environments Prove Alignment Is Hard(4 posts)→

Original post →

More from Research

Research channel →