Zitronbench: user challenges critic Ed Zitron to name a task SOTA LLMs fail

iruletheworldmo · x · 2026-09-06

An X user has publicly challenged AI critic Ed Zitron to name a simple task that current SOTA large language models get wrong. The challenge, dubbed zitronbench, invites anyone in the replies to one-shot Zitron's task with proof it was LLM-only. It's a public wager against the recurring claim that LLMs fail at even the most basic tasks.

Original post →

More from Fun

Fun channel →