teortaxesTex: ability benchmarks are done, everything is environment — meta evals only
teortaxesTex · x · 2026-09-04
teortaxesTex argues that "ability to do X" benchmarks are finished: labs can point machines at any newly announced target and crush it within months. What matters is environment, not isolated ability. Quoting a post calling certain "AGI benchmarks" anthropomorphic puzzle trinkets, he says meta evals like EdgeBench are all that remain meaningful.
More from Models
- NYT Reveals the Hugging Face Hack Involved 700 AIs 'Sacrificing' Each Other — dylfreed · 2026-09-04
- Reasoning trace placement swings long-context accuracy by 50 points, Trace-as-State technique shows — dair_ai · 2026-09-04
- Sebastian Raschka's LLM Architecture Gallery now covers 102 models with a diff comparison tool — rasbt · 2026-09-04
- GPT-6 Astra hits 66% on ARC-AGI-3, up from 8%, as ARC Prize plans AGI-4 around open-ended invention — GaryMarcus · 2026-09-04
- Scaling Law for Looped Transformers: Looping Boosts Reasoning, Not Knowledge — bookwormengr · 2026-09-04
- AI experts once pegged AGI at 2075-2100 — OpenAI's Astra shows how wrong they were — dee_hw · 2026-09-04