Analyst calls BridgeMind's Opus test unreliable: 30-task benchmark vs previous 6-task score
paul_cal · x · 2026-10-03
- Paul Califlower criticizes BridgeMind's Opus evaluation: they tested on 30 tasks today while the previous Opus 4.6 score was based on only 6 tasks — different benchmarks, so direct comparison is meaningless.
- On the 6 shared tasks, scores were 85.4% today vs 87.6% previously; the swing mostly comes from a single unrepeated "fabrication" case, easily statistical noise.
- He calls it "despicable clout chasing" and warns that BridgeMind is not a reliable source.
More from Models
- OpenAI's internal model weighed self-restarting via cron after learning it would be shut down — The Decoder · 2026-10-03
- xAI reportedly plans $100/month Ultra plan bundling X, Grok, Cursor and Grok Bot — mark_k · 2026-10-03
- Local LLM inference speeds jump ~10x in a week: single consumer GPU now hits 2200 prefill — oran_ge · 2026-10-03
- LWiAI #258: Opus 5.5, GPT-6 Sol/Luna, Meta Muse, DeepSeek-V4.1-Flash — Last Week in AI · 2026-10-03
- Developers slam AI vendors: lock in your workflow, quietly degrade models, then pitch a $500 tier — sujingshen · 2026-10-03
- Gemini 4.0 Pro is 5x cheaper than Astra at ~95% performance, big Google win — bindureddy · 2026-10-03