Developer argues 100-1000 tps LLM speed is pointless unless rewriting legacy code to Rust in one pass

ssh4net · x · 2026-10-01

A developer questioned the value of pushing LLM inference to 100-1000 tokens per second: unless you're rewriting a massive legacy project to Rust in a single pass, the extra speed doesn't matter much.

He adds that in a systematic development workflow, 99% of the time goes to project builds and running extensive tests, so model output speed is not the bottleneck.

Original post →

More from coding & agent

coding & agent channel →