Alex Zhang: design the language model's shape around agents, with recurrent memory for old history
CatAstro_Piyush · x · 2026-10-01
Stanford researcher Alex L. Zhang published a blog post, "Language Model 'Shape'", arguing the input/output shape of LLMs has been static since ChatGPT, and all agent work fits a harness around the autoregressive decoder-only Transformer. He asks whether it's worth flipping that: change the model's shape to fit the harness.
Key points:
- Modifying the harness is cheap, hence dominant — but amortized over scaling, changing the model shape may cost less
- Modern architecture research only tweaks layer details because decoder-only models are powerful, architecture bets are worth millions, and next-token prediction is a natural objective
- His concrete idea: recurrent memory for old history + dense attention for recent context, reducing reliance on manual context compaction
More from coding & agent
- Ask AI to record itself trying your product for the first time — lucasmeijer · 2026-10-01
- Box CEO Aaron Levie: every company will need an army of forward deployed engineers for AI transformation — rohanpaul_ai · 2026-10-01
- Meta-reasoning harness hits 71.5% on ProgramBench, beating Codex by 13.5 points — anirudhg9119 · 2026-10-01
- Moda launches Slack agent for querying alerts, user intent, and traces — KlausCodes · 2026-10-01
- Cognition Becomes First CoreWeave Vera Rubin NVL72 Customer, Sees 4.8X SWE-2 Throughput Boost — altryne · 2026-10-01
- LlamaIndex Draws ~600 in SF for Agent Document Processing Events — llama_index · 2026-10-01