Matthew Berman: OpenAI pauses testing after Hugging Face hack as benchmarks race ahead
Matthew Berman · youtube · 2026-09-11
In this video, Matthew Berman connects a wave of recent milestones and warning signs: OpenAI's Paperbench evaluation, Anthropic's Claude Opus 4.5 launch, DeepMind's AlphaEvolve coding agent, METR's time-horizons research, and OpenAI pausing development and testing after the Hugging Face hack. He also revisits the Pause Giant AI Experiments letter and the 'Pacing the Frontier' essay, arguing that as capability metrics climb fast, frontier labs are now hitting the brakes over safety incidents — and frontier pacing is becoming the industry's central question.
More from Models
- ChatGPT monthly active users top 1.06 billion in August, fourth straight record month — FinanceYF5 · 2026-09-11
- PuzzleMask: Plain-Prose Attack Bypasses All 4 Tested LLM Gatekeepers at 100% — TechNadu · 2026-09-11
- OpenAI Codex may issue another usage reset this weekend, says Codex lead resets happen — umesh_ai · 2026-09-11
- OpenAI Reportedly Pointing Its Navier–Stokes Model at Riemann and P vs NP — 141_1337 · 2026-09-11
- Benchmark author says OpenRouter unreliably honors Meta Muse effort levels, EU payments broken — PawelHuryn · 2026-09-11
- User burns $200 of Codex credits in one agent turn — 4,700 of 5,000 credits, task unfinished — RileyRalmuto · 2026-09-11