Benchmark where LLMs scored under 1% on Fiverr tasks quietly abandoned, sparking skepticism

BLUECOW009 · x · 2026-09-05

BLUECOW009 quotes CryptoCyberia's skepticism: the benchmark that tested LLMs on real Fiverr tasks — where every model scored under 1% — has quietly been abandoned. "The fact that they abandoned this benchmark really makes you think..." BLUECOW009 adds: "never make a move when your opponent is making a mistake," implying vendors and eval teams are dodging exposure of models' weaknesses on messy real-world tasks.

Original post →

More from Fun

Fun channel →