Smaller LLMs aren't always cheaper: the hidden costs nobody measures

zeuslac · reddit · 2026-09-21

A cheaper model per call can make the whole workflow more expensive. On easy cases savings are real; on hard cases reviewers check the weak model's output against sources and redo the work themselves, making those cases slower than pre-AI, while errors that slip through cost far more to fix later.

Two implications: difficulty-based routing is only worth it given the price gap — a single price drop can erase its reason to exist, so keep it cheap to unwind; and a router can be financially better while pushing overall error rates above the manual baseline.

The fix: one record per task tying model usage, review time, and later corrections together, so you measure the full cost of a completed task, not the cost of a call.

Related event: Smaller Models Can Cost More: Hidden Review and Rework Expenses(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →