PSA: tok/s comparisons across tokenizers are unit illusions — convert to chars/token first
challis88ocarina · reddit · 2026-09-29
A PSA explains why comparing tokens/sec across model families is often an illusion: a token is a slice from each model's private vocabulary, so tok/s measures 'steps per minute,' and models have different stride lengths.
Key rules:
- Same tokenizer = fair comparison: two llama.cpp servers on the same model family can be compared directly.
- Different tokenizers = different units: convert to real speed — both servers reporting 40 tok/s may deliver 131 vs 193 chars/s, a 1.48x real difference. Choppy tokenizers inflate headline numbers.
- The conversion factor is content-dependent: maybe 0.68 chars/token on prose, 0.74 on JSON — measure on representative workloads, don't trust vendor marketing.
- One real, one illusion: efficient tokenizers genuinely reduce prefill compute and TTFT (a real cost win); cross-family decode-rate comparisons say nothing about who finishes first.
Bottom line: any cross-server speed claim should come with chars/token (or words/sec) on a representative workload.
More from Infra
- Weaviate Podcast: split database duties between Postgres for structured data and vector search — CShorten30 · 2026-09-29
- Fireworks Launches FireRouter: 98.1% of Opus Accuracy at 57% Lower Coding Cost — nicolechirps · 2026-09-29
- Data center developer offers $10,000 checks to 4,500 PA households if 1,300-acre facility approved — Polymarket · 2026-09-29
- SpaceX Building Terafab at Full Speed: Program Launched in March, Concrete Already Pouring — scaling01 · 2026-09-29
- Grok 4.7 hits Amazon Bedrock: 500K context, 2x output tokens for the gains — AWS ML Blog · 2026-09-29
- Google Is Sending an AI Data Center to Space: Satellite With Onboard Compute Launches Next Week — bratton · 2026-09-29