GPT-6 Beats All 48 Levels of Neal.fun's 'I'm Not a Robot' Game as CAPTCHAs Crumble
APPSO · wechat · 2026-09-08
Neal Agarwal's 'I'm Not a Robot' browser game on neal.fun has 48 increasingly absurd levels (find Wally, draw a perfect circle, break up with your 'girlfriend'), and GPT-6 has now cleared them all — a vivid sign that visual CAPTCHAs are losing the arms race.
- The data: A 2023 USENIX Security study found bots solve mainstream CAPTCHAs at 85–100% accuracy, beating humans' 50–85%; a 2024 ETH Zurich PhD project showed Google's checkbox is trivially defeated by targeted models.
- A short history: CMU's warped-text Gimpy (2000, used by Yahoo) was 30% cracked within two years; Google's reCAPTCHA checkbox (2014) was bypassed at 70% accuracy by 2016.
- The pivot: The industry is abandoning challenges for invisible scoring — Cloudflare Turnstile, reCAPTCHA v3's 0–1 human-probability score — and Wikipedia switched to behavior-based hCaptcha this year, challenging only 0.1% of users visually.
- Context: Cloudflare estimates bot traffic rose from 30% to 60% of the web in a year, and will block mixed-purpose AI crawlers on ad-supported pages by default from Sep 15, 2026.
- The deeper problem: AI agents booking tickets or research on your behalf are technically bots with legitimate intent; verifying humanity doesn't verify intent. Product design needs revocable authorization and anomaly cutoffs instead. Google is experimenting with camera-based gesture CAPTCHAs, raising new privacy and false-positive concerns.
GPT-6 also scores 92.7% on ScreenSpot-Pro (tool-free), vs 76.9% for GPT-5.6 Sol — moving from 'knowing which element to click' to actually locating and operating it.
Related event: GPT-6 Astra Clears All 48 Levels of the "I'm Not a Robot" CAPTCHA Game(6 posts)→
More from Models
- LinkedIn Is LLMs' #2 Cited Site — Your Posts Are Becoming AI Knowledge — jaindl · 2026-09-08
- Benchmaxxing exposed: fresh benchmark drops model scores from 89% to 19% — Able-Line2683 · 2026-09-08
- Dev prefers Fable 5.1/5.6 Sol for planning and coding, eyes Grok 4.7 next — iannuttall · 2026-09-08
- Minimax H3 speedup modes strip realism, full steps needed for production — Choiced_Gamer · 2026-09-08
- User burns 40% of weekly GPT 6 Astra limit in 11 hours, complaining quotas drain fast — iScienceLuvr · 2026-09-08
- AI dictation still lags human voiceover: even SOTA ElevenLabs tires on long content — ethanniser · 2026-09-08