Paper: 3 counting examples get gpt-3.5-turbo to 99% on strawberry, showing thinking is computational

ctjlewis · x · 2026-09-28

Author ctjlewis explains the core idea of his paper: thinking is computational. Today's natural-language "reasoning" in LLMs is structurally the same thing as the pure-symbol computation in his paper, just at the level of ideas instead of symbols.

The headline experiment: adding just 3 worked counting examples to the prompt gets gpt-3.5-turbo to roughly 99% accuracy on the classic "how many r's in strawberry" question — but only because the model is forced to do explicit intermediate work counting the letters. The takeaway is that such failures reflect missing intermediate computation, not a hard capability ceiling.

He also notes recent work pretraining models on randomized Universal Turing Machines before RL, calling it solid work the field is starting to appreciate.

Related event: Running cellular automata on an 'LLM computer' to argue thought is computation(2 posts)→

Original post →

More from Models

Models channel →