If early autocomplete-era LLMs could code, do we even need post-training?

menhguin · x · 2026-09-20

banteg resurfaces the 2019 Deep TabNine announcement—a GPT-2 code completion model trained on 2 million GitHub files—to revisit an old question: is programming more about language or logic?

His observation: early LLMs were essentially "linguistic autocomplete," yet already handled programming decently despite their simplicity, suggesting coding is surprisingly friendly to pure language modeling. That leads to his provocative question: if pretraining alone gets you this far, maybe we don't need post-training at all.

Original post →

More from AGI Musings

AGI Musings channel →