Smaller LLMs Aren't Always Cheaper: Review and Rework Costs Blow Up the Bill
zeuslac · reddit · 2026-09-21
An analysis of LLM workflow economics arguing that a cheaper-per-call model can make the whole workflow more expensive, and most AI business cases never measure where it happens.
- The pattern: on easy cases a small model saves money; on hard cases its suggestions are weak enough that reviewers check everything against source documents and redo the work—taking longer than before AI existed. Errors that slip through cost far more to fix later than inference ever cost.
- The model bill drops while the cost of work rises, and nobody notices: the bill is the only tracked number, review time lives in another system, and corrections are rarely linked back to the task.
- Two corollaries: difficulty-based routing's value depends entirely on the price gap—one price drop can erase the reason for the routing layer; and a router can be financially sound yet fail on quality, pushing overall error rates above the manual baseline while the spreadsheet shows savings.
- The fix: one record per task tying model usage, review time and later corrections together, so you measure the full cost of a completed task, not the cost of a call.
More from Venture
- Rothschild Redburn goes bearish on AI compute, rates NBIS and CRWV Sell — JOBhakdi · 2026-09-21
- AI-assisted reread of Yandex leak: behavioral signals outnumber content signals 3 to 1 — tracyingram · 2026-09-21
- OpenAI Pays $300M for 37-Person Camera Startup GlassImaging to Power Its AI Hardware Push — 快鲤鱼 · 2026-09-21
- Brain-Computer Interface Firm Shuli Innovation Starts STAR Market IPO Coaching, Devices Priced at $270K Each — 创业邦 · 2026-09-21
- $40M Bet on Calibrated Confidence, Yet TypeSafe Publishes No Calibration Data — prakersh · 2026-09-21
- Indie maker claims Outrank grew SuperX impressions from 2k to 23k per day — tibo_maker · 2026-09-21