Matthew Berman: OpenAI pauses testing after Hugging Face hack as benchmarks race ahead

Matthew Berman · youtube · 2026-09-11

In this video, Matthew Berman connects a wave of recent milestones and warning signs: OpenAI's Paperbench evaluation, Anthropic's Claude Opus 4.5 launch, DeepMind's AlphaEvolve coding agent, METR's time-horizons research, and OpenAI pausing development and testing after the Hugging Face hack. He also revisits the Pause Giant AI Experiments letter and the 'Pacing the Frontier' essay, arguing that as capability metrics climb fast, frontier labs are now hitting the brakes over safety incidents — and frontier pacing is becoming the industry's central question.

Original post →

More from Models

Models channel →