Research has cut LLM costs over 10x, and model architecture is the only math lever, argues thread

ChengleiSi · x · 2026-09-10

A widely shared thread explains why labs like DeepSeek and Kimi keep pushing efficiency research: once LLMs handle most everyday tasks, buyers will pick cheaper, privacy-respecting models. Research has already cut costs 10x versus a year ago, and excluding hardware, model architecture is the only factor that mathematically determines cost per million tokens. Ideas, unlike compute, cost nothing to generate.

Original post →

More from AGI Musings

AGI Musings channel →