Researchers debate whether prompt shaping or token-level input is what actually drives the gains

a1zhang · x · 2026-07-23

A technical reply arguing that the experiment is not proving raw LLMs “need token-by-token input,” but rather showing that a harness with prompt shaping / task decomposition can materially improve performance.

The author says it is obvious an individual LLM will do well on prompts it saw during training, and says the useful question is whether the harness’s decomposition is necessary. Their view: the experiments show it is valuable, though not necessarily strictly necessary. They also frame the comparison against a base Transformer as a test of whether the model generalizes to structurally similar inputs, or to the same task at different context lengths. The thread is really about what the experiment does and does not demonstrate.

Related event: LLM Generalization Debate: Intrinsic Ability or External Harness?(11 posts)→

Original post →

More from coding & agent

coding & agent channel →