A private eval with a 0% completion rate for 3 years: no AI model can identify this flag
generativist · x · 2026-09-24
Developer teejm shares a private eval he has run for 3 years with a 0% completion rate: identifying a picture of a flag. The flag exists on Wikipedia and Google Images and is seen by millions of people yearly, yet no AI model on the planet can name it. He notes the new Astra model gets "really, really close" — the closest any model has come. The post highlights a persistent long-tail blind spot in multimodal recognition.
More from Fun
- Joke: One Week to Squeeze Opus 5.5 Before It Gets Nerfed to Haiku 3.5 — vasuman · 2026-09-24
- Indie dev's Mid-Autumn interactive site goes viral after OpenAI's like — dotey · 2026-09-24
- CC Switch dev found stars flatlined after moving downloads off GitHub — tinyfool · 2026-09-24
- GPT-6 Sol Max is the new Luna Max: the model you limp along on until limits reset — ___Patrice___ · 2026-09-24
- Three.js veteran jokes he can't tell Three.js apart from physical reality — nptacek · 2026-09-24
- AI meme flips the script: "We need to slow down Human Intelligence!" — telestitchtv · 2026-09-24