Developer argues 100-1000 tps LLM speed is pointless unless rewriting legacy code to Rust in one pass
ssh4net · x · 2026-10-01
A developer questioned the value of pushing LLM inference to 100-1000 tokens per second: unless you're rewriting a massive legacy project to Rust in a single pass, the extra speed doesn't matter much.
He adds that in a systematic development workflow, 99% of the time goes to project builds and running extensive tests, so model output speed is not the bottleneck.
More from coding & agent
- Hybris MCP Server lets AI assistants manage SAP Commerce Cloud instances — modelcontextprotocol · 2026-10-01
- cua-speedrun: CMU benchmark shows 4.4x speed gap between equal-scoring computer-use agents — arankomatsuzaki · 2026-10-01
- M-Anchor: a deterministic, zero-LLM Python gate that blocks unsupported LLM record updates — Informal-Winter-3190 · 2026-10-01
- Dev Burns Through Claude Weekly Quota in Under Two Days on Opus 4.5 Alone — yihui_indie · 2026-10-01
- Split coding workflow: Opus 5.5 xhigh for planning, Sol 6.1 high for the bulk of coding — haider1 · 2026-10-01
- FV expert: AI folks' view of formal verification is 15 years out of date — tianyin_xu · 2026-10-01