Model slips on a science knowledge test, author calls it acceptable
felpix_ · x · 2026-09-08
The author shared a screenshot comparing a model's performance on a science knowledge test, noting an unfortunate drop in scores on that benchmark. In a follow-up reply he judged the result acceptable overall. Specific model names and scores are only visible in the attached image.
More from Models
- Don't hand off to a cheap model: Luna fails as a subagent for Astra on tasks needing real understanding — brandon_galang · 2026-09-08
- Real-world test punctures 'AGI has arrived' hype: Astra botches a simple website task — GaryMarcus · 2026-09-08
- Rumor: GPT-6 Astra reportedly launched, OpenAI president hails 'the AGI era' — aftahi_ai · 2026-09-08
- Unverified claim: GPT-6 Astra scores 91.8% on SpatialBench, beating 80% human baseline — paigeinsf · 2026-09-08
- Fable 5.1 ships with worse benchmarks, but devs say it's better at real engineering — _arohan_ · 2026-09-08
- Mollick: trillion-dollar AI industry still judged by crude evals despite better methods existing — emollick · 2026-09-08