AA-Omniscience benchmark: 6,000 fact questions show top models still err ~25% without search
randal_olson · x · 2026-09-30
Randall Olson shared Artificial Analysis' AA-Omniscience data: 6,000 questions about specific facts across 42 work topics, answered without search. Wrong answers are measured as a share of answers given (declined questions excluded).
Headline finding: even the best models get roughly 1 in 4 answers wrong without search — if a fact matters, have the model look it up.
He also released evident-charts, an open source agent skill for Claude Code, Codex, Cursor and more that teaches AI coding agents to make clear, honest charts: it checks data first (totals mixed with parts, duplicate rows, placeholder codes, preliminary months) and critiques and rebuilds bad drafts — e.g., raw counts on a log scale, the EU total ranked among its own members, meaningless rainbow palettes.
Related event: Top AI Models Get About a Quarter of Factual Questions Wrong Without Search(2 posts)→
More from Models
- Defending Astra's Caution: An Agent That Stops to Ask Is Doing It Right — brandon_galang · 2026-10-01
- Dev complains Opus 5.5 burned 80% of quota in 3 days, switching models — MaziyarPanahi · 2026-10-01
- Liquid AI's LongevityBench: compact LFMs beat frontier models on several aging tasks — JosephJacks_ · 2026-10-01
- GPT-6.1 Sol hands-on: efficient but slow, and subscription changes worry users — kimmonismus · 2026-09-30
- GPT 6.1 Astra ultra code mode builds a three.js ocean scene — OpenAIDevs · 2026-09-30
- Analysis: OpenAI's DOTS may be the answer to ChatGPT's growing product sprawl — mark_k · 2026-09-30