Community inference contest hits 1900 tps prefill, 2.5x faster than omlx baseline
HankYeomans · x · 2026-10-02
- Results: David Tai announced the inference optimization contest ended with up to 450% gains — 1900 tps prefill and 433 tps decode, roughly 2.5x the omlx reference benchmarks from a week earlier.
- How: Beyond kernel upgrades to the model runner, contestants tuned dflash2 to de-weight previous tokens when drafting; Bonsai 2's freed compute bandwidth surfaced many new techniques worth open-sourcing.
- Next: The contests return next week while organizers improve infrastructure and address resubmission reward gaming.
More from Infra
- The Neocloud Reality Check: Why Your Next AI Project May Skip the Big Clouds Entirely — DavidLinthicum · 2026-10-02
- Ex-SWE Turned Inference Engineer: Ollama Is Never Optimal — A Bottleneck Guide — Postmodern_Plunger · 2026-10-02
- A 3-step guide to open models: picking, hosting locally or via OpenRouter — every · 2026-10-02
- San Antonio district hosts a dozen data centers as industry camouflages them in woods — The Verge AI · 2026-10-02
- Need a GPU fast? Self-serve clouds like RunPod offer instant spin-ups without contracts — DavidLinthicum · 2026-10-02
- Amazon writes 3,000-word blog warning communities not to block data centers — The Verge AI · 2026-10-02