Andon Labs Replay: GPT-4o Makes Termination Decisions Far Less Often Than Frontier Models

tokenbender · x · 2026-08-16

tokenbender argues that terra doesn't feel like a smaller sol — the post-training gap between the two looks substantial, showing up as a big personality difference. Andon Labs replied "you're onto something": when they replayed the scenario, GPT-4o made the termination decision far less frequently than frontier models did.

Original post →

More from Models

Models channel →