One Model, Three Skills: Programmatic Use, Chat, and Test-Taking Diverge
lateinteraction · x · 2026-09-19
The author argues that debates about LLM capability often conflate three distinct problems: the abstractions and training that make a model good at programmatic use are very different from those that make it good at user-facing interactions, which in turn differ from those that make it good at test taking.
This responds to Nabeel Quazy's observation that LLMs feel like geniuses in fun personal projects yet behave stupidly when deployed in messy real-world environments — precisely because the two scenarios rely on different training paths.
More from Models
- Noam Brown: GPT-6 Astra does have observable chain of thought, calls it fragile — burny_tech · 2026-09-19
- Jev's popularity signals the AI crowd is open to models beyond LLMs — BLUECOW009 · 2026-09-19
- QuixiAI picks gemma-4-26B-A4B-it as base model for OpenJev — QuixiAI · 2026-09-19
- Open-weight Jev replica based on Qwen3.8 27B with 265k context drops tomorrow — TheZachMueller · 2026-09-19
- Jev beats a Sonnet 5-powered retriever on accuracy at a fraction of the cost — IanArawjo · 2026-09-19
- Alibaba open-sources medical AI model that detects cancer and nearly 150 conditions — giveen · 2026-09-19