Smaller LLMs aren't always cheaper: the hidden costs nobody measures
zeuslac · reddit · 2026-09-21
A cheaper model per call can make the whole workflow more expensive. On easy cases savings are real; on hard cases reviewers check the weak model's output against sources and redo the work themselves, making those cases slower than pre-AI, while errors that slip through cost far more to fix later.
Two implications: difficulty-based routing is only worth it given the price gap — a single price drop can erase its reason to exist, so keep it cheap to unwind; and a router can be financially better while pushing overall error rates above the manual baseline.
The fix: one record per task tying model usage, review time, and later corrections together, so you measure the full cost of a completed task, not the cost of a call.
Related event: Smaller Models Can Cost More: Hidden Review and Rework Expenses(3 posts)→
More from AGI Musings
- A decade later, he admits: dismissing AI safety because of its advocates was a mistake — austinc3301 · 2026-09-21
- Recap: Salesforce-Anthropic Claudeforce adds 37 sales skills; ByteDance five-stage self-improvement roadmap — thione · 2026-09-21
- Chinese researchers outline five-stage roadmap for AI self-improvement without humans — thione · 2026-09-21
- No AI-driven developer unemployment in the data: US dev share and salaries up vs 2021 — lemire · 2026-09-21
- Trask: The 'artists' in AI debates are small business owners, not superstars — iamtrask · 2026-09-21
- a16z: AI doubled app supply but downloads stayed flat — App-Slop is the story — a16z · 2026-09-21