Jev fails counting r's in strawberry, aces it 168/168 when given letters

BLUECOW009 · x · 2026-09-22

redp314 ran the classic strawberry test on @typesafeai's Jev: it failed like every LLM—47% said 3 r's, 47% said 2, and it undercounted doubled letters on 70% of 168 test words.

But when given the letters as a list instead of the word, the same model answered 168/168 correctly in 260ms—pointing to tokenization, not reasoning, as the culprit. Quipped as a "skill issue."

Original post →

More from Models

Models channel →