GPU contest: batched compact-Householder QR kernel achieves 232x speedup
petrusenko_max · x · 2026-08-15
In GPU Mode's contest, a batched compact-Householder QR kernel achieved a 232x speedup over baseline, ranking 12th of 183. Over 14 days, 1500+ submissions enabled iterative improvements using feedback on runtime across matrix sizes.
More from Infra
- Google floats many TPU RFPs, takes more wafers directly to TSMC each generation — BenBajarin · 2026-08-16
- Baseten adds Day 0 support for finetuning Qwen3.8-27B — baseten · 2026-08-16
- llama.cpp adds support for Moonshot Kimi-K3 text model — pmttyji · 2026-08-15
- Self-Hosting Qwen3.8-27B NVFP4 on SGLang Hits 200+ Tokens/Sec — gnukeith · 2026-08-15
- AMD's internal silicon photonics group and Ayar Labs collaboration may surpass MediaTek — zephyr_z9 · 2026-08-15
- Nvidia Cuts OpenAI Data Center Investment to $120B; Anthropic Revenue Surges — The Decoder · 2026-08-15