Jev fails counting r's in strawberry, aces it 168/168 when given letters
BLUECOW009 · x · 2026-09-22
redp314 ran the classic strawberry test on @typesafeai's Jev: it failed like every LLM—47% said 3 r's, 47% said 2, and it undercounted doubled letters on 70% of 168 test words.
But when given the letters as a list instead of the word, the same model answered 168/168 correctly in 260ms—pointing to tokenization, not reasoning, as the culprit. Quipped as a "skill issue."
More from Models
- Vals AI: Grok 4.7 drops to #24 on Vals Index, down 5 points from Grok 4.6 — zacharynado · 2026-09-22
- Grok 4.7 example: three-year financial analysis exposes currency-masked growth stall — ArtificialAnlys · 2026-09-22
- AA example: Grok 4.7 independently runs valuation chain and flags divergence from deal partner — ArtificialAnlys · 2026-09-22
- Grok 4.7 ranks just behind Anthropic's Opus 5 on AA-Briefcase at ~50% of the cost per task — ArtificialAnlys · 2026-09-22
- Replicating ExploitBench Would Cost ~$59.3M in API Fees, Security Researcher Estimates — OwariDa · 2026-09-22
- Xiaomi Releases Small Qwen 3.5 9B Distill SFT'd on MiMo Data, Plus RL Environments — teortaxesTex · 2026-09-22