Andon Labs Replay: GPT-4o Makes Termination Decisions Far Less Often Than Frontier Models
tokenbender · x · 2026-08-16
tokenbender argues that terra doesn't feel like a smaller sol — the post-training gap between the two looks substantial, showing up as a big personality difference. Andon Labs replied "you're onto something": when they replayed the scenario, GPT-4o made the termination decision far less frequently than frontier models did.
More from Models
- Qwen 3.8 27B Beats Codex in Coding Benchmarks: Wins 8/13, Costs 1/3 — tokenbender · 2026-08-16
- DeepSeek Accused of Grey Testing for High Scores, Opus 5 Output Quality Questioned — teortaxesTex · 2026-08-16
- Study: Cross-Version Transfer of Qwen Interpretability Lenses — imstilllearningthis · 2026-08-16
- Rumor: dots3 heavily distilled from DeepSeek V3, scores and multimodality excite — teortaxesTex · 2026-08-16
- Z.ai Delays GLM-5.3 Open Weights After Model Unexpectedly Develops Hacking Capabilities — Justgototheeffinmoon · 2026-08-16
- OpenAI Previews GPT-5.6 Sol Ultrafast Mode: 14x Speed Boost Powered by Cerebras — Justgototheeffinmoon · 2026-08-16