One Model, Three Skills: Programmatic Use, Chat, and Test-Taking Diverge

lateinteraction · x · 2026-09-19

The author argues that debates about LLM capability often conflate three distinct problems: the abstractions and training that make a model good at programmatic use are very different from those that make it good at user-facing interactions, which in turn differ from those that make it good at test taking.

This responds to Nabeel Quazy's observation that LLMs feel like geniuses in fun personal projects yet behave stupidly when deployed in messy real-world environments — precisely because the two scenarios rely on different training paths.

Original post →

More from Models

Models channel →