Alignment researcher's 2022 "moment of terror": GPT-4 done training, mentors quit OpenAI
JacquesThibs · x · 2026-09-14
AI safety researcher JacquesThibs recounts his August 2022 "moment of terror": days of backcasting alignment on a whiteboard kept pointing to deception and Goodharting derailing everything; his mentor, spooked right after GPT-4 finished training, quit OpenAI believing shorter timelines left nothing to do from inside; another former mentor, aware of the completed training and an upcoming capability jump, became convinced of a 2026 timeline. Notably confirms GPT-4 had finished training months before release.
More from AGI Musings
- Nina Schick: public distrusts regulators as much as AI labs — nobody can 'pace' AI — NinaDSchick · 2026-09-14
- Jensen Huang: next 2 decades of progress may exceed all of history combined — rohanpaul_ai · 2026-09-14
- Researcher: I'm less worried about AI x-risk than in 2022 — time to retire p(doom) — soumitrashukla9 · 2026-09-14
- Security veteran to AI labs: capability isn't risk, cyber evals lack real-world threat modeling — HackingLZ · 2026-09-14
- GPT-4o psychosis snippets eerily resemble SCP wiki stories, likely in training data — code_star · 2026-09-14
- Vals AI: frontier labs shouldn't grade their own frontier; models may match researchers by Aug 2027 — JenniferHli · 2026-09-14