teortaxesTex: ability benchmarks are done, everything is environment — meta evals only

teortaxesTex · x · 2026-09-04

teortaxesTex argues that "ability to do X" benchmarks are finished: labs can point machines at any newly announced target and crush it within months. What matters is environment, not isolated ability. Quoting a post calling certain "AGI benchmarks" anthropomorphic puzzle trinkets, he says meta evals like EdgeBench are all that remain meaningful.

Original post →

More from Models

Models channel →