Multi-LLM social worlds are 'large': models may learn heuristics instead of grader-modeling
1a3orn · x · 2026-09-11
Continuing his discussion with jdpressman, 1a3orn argues multi-LLM training is probably complex enough that rather than fully modeling the grader, models might instead learn something like desired heuristics, since social worlds are 'large' worlds.
The author admits he isn't really sure, but the point echoes his earlier argument that environment design determines whether models solve problems or obsess over the grader.
Related event: Researcher Questions Whether Bad RL Environments Prove Alignment Is Hard(4 posts)→
More from Research
- RD-Forget: reversible, query-dependent forgetting for agent memory — MaryamMiradi · 2026-09-11
- Science Advances editor: no author has ever disclosed AI use despite policy — TuhinChakr · 2026-09-11
- YOCO explained: one shared KV cache reused across the model's second half — stochasticchasm · 2026-09-11
- PARSER: parallel chunk subagents with an RL-trained lead agent for long-context QA — omarsar0 · 2026-09-11
- Year-long study: heavier AI companion engagement predicts lower well-being — dhadfieldmenell · 2026-09-11
- DeepMind launches AlphaGenome Atlas, a 1TB navigable map of human DNA — neil_chilson · 2026-09-11