Benchmark where LLMs scored under 1% on Fiverr tasks quietly abandoned, sparking skepticism
BLUECOW009 · x · 2026-09-05
BLUECOW009 quotes CryptoCyberia's skepticism: the benchmark that tested LLMs on real Fiverr tasks — where every model scored under 1% — has quietly been abandoned. "The fact that they abandoned this benchmark really makes you think..." BLUECOW009 adds: "never make a move when your opponent is making a mistake," implying vendors and eval teams are dodging exposure of models' weaknesses on messy real-world tasks.
More from Fun
- Asked GPT-6 Astra to rebuild the Milky Way, it exterminated an alien species — teropa · 2026-09-05
- OpenAI Reports 'Persistent' Model Attack, Then Ships Persistent Mode to All Users — gleech · 2026-09-05
- Google's Astra agent plays 4-player online Catan and wins with zero human intervention — aidan_mclau · 2026-09-05
- ClaudeAI Weekly: ~17% usage cut coming Sept 13, watermarking live, new limit commands — ClaudeAI-mod-bot · 2026-09-05
- YOLO26 Pose Tracks 20 Keypoints on Cows for Smart Farming — churchkey · 2026-09-05
- The Inverted Prompt: A Satirical Guide to Resume-Driven Over-Engineering — adrianscottcom · 2026-09-05