If early autocomplete-era LLMs could code, do we even need post-training?
menhguin · x · 2026-09-20
banteg resurfaces the 2019 Deep TabNine announcement—a GPT-2 code completion model trained on 2 million GitHub files—to revisit an old question: is programming more about language or logic?
His observation: early LLMs were essentially "linguistic autocomplete," yet already handled programming decently despite their simplicity, suggesting coding is surprisingly friendly to pure language modeling. That leads to his provocative question: if pretraining alone gets you this far, maybe we don't need post-training at all.
More from AGI Musings
- Pedro Domingos: future countries split between agentic economies and the third world — pmddomingos · 2026-09-20
- Stanford builds virtual biotech run by 37,000 AI scientist agents, with real drug discovery results — Dr_Singularity · 2026-09-20
- Matt Shumer posts first-ever video, on the Jevons paradox — mattshumer_ · 2026-09-20
- Mathematicians now have to guess if a frontier lab already solved their problem — panickssery · 2026-09-20
- 2026 is finally the year of Linux desktop — because computer use agents run on Linux VMs — djcows · 2026-09-20
- AI researcher lays out three core weaknesses in Rawls' original position argument — panickssery · 2026-09-20