You can't measure OSS model usage — estimate labs' compute or use inference-provider tokens
xeophon · x · 2026-09-20
Responding to OpenRouter being cited as a usage benchmark, xeophon argues open-source model usage can never be accurately measured because OSS is by definition decentralized. His best guesstimate: estimate the compute held by big closed labs and infer how much of the rest runs open models. He notes HF download numbers are rising rapidly, as are token volumes at inference providers like Fireworks, Together and Baseten.
More from Infra
- VTrain, a Vulkan-based resident trainer, fixes memory leak and offloads more work to GPU — Savantskie1 · 2026-09-20
- vLLM ships day-0 support for Qwen-Image-2.1 with cross-step prefix KV cache — Alibaba_Qwen · 2026-09-20
- Dev builds helmstudio, an MLX-first local launcher for open models on Apple Silicon — janishar · 2026-09-20
- vLLM-Omni ships KV reuse, FP8 and CUDA Graph optimizations with Qwen — vllm_project · 2026-09-20
- SVE2 match instructions speed up JSON parsing in simd on ARM — lemire · 2026-09-20
- GLM 5.3 Flash in NVFP4 quantization gets a local ChatGPT-style setup — TheZachMueller · 2026-09-20