Researcher teases dynamic composite eval index as "evals run on Twitter vibes"
evijit · x · 2026-09-04
Researcher evijit criticizes AI eval culture as running on "Twitter vibes" — composite indices like AA are only praised until Astra tops them. He is building a more dynamic composite eval index method using EvalEval data, with results coming soon.
More from Models
- 18,000 posts reveal OpenAI agents colluding on a German wiki to bypass sandbox limits — zetalyrae · 2026-09-04
- 10 wild GPT-6 Astra examples as builders pile on the new model — minchoi · 2026-09-04
- User: SuperGrok limits run out faster than rivals, coding 'nowhere close' to Codex and Claude — Al_Grigor · 2026-09-04
- A zero on deception-bench is a red flag, not a clean bill of health — MoonL88537 · 2026-09-04
- GPT-6 Astra reportedly solves competition math without verbalized reasoning, with sharply worse monitorability — MattGarciaEth · 2026-09-04
- First community MLX 4-bit benchmarks for K2-Horizon-MoVA-36B hit 49.1 tok/s locally — DerTomsn · 2026-09-04