QR Kernel Runs 280x Faster on B200
dejavucoder · x · 2026-07-13
[Forwarded Content] The author secured 5th place in the GPU MODE qrv2 competition with a batched QR decomposition kernel that is roughly 280x faster than torch.geqrf on the B200.
They also mentioned their autoresearch swarm ran about 2000 experiments in two weeks. The full writeup covers:
- Householder QR algorithm concepts
- Kernel optimization techniques
- A glimpse into "agentic kernel research" in 2026
More from Infra
- How to build a PostgreSQL-backed semantic search pipeline with pgvector and Ollama — KhuyenTran16 · 2026-07-21
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- Milled from Solid Aluminum: AI Rig Multi-GPU Case for Local Compute — dee_hw · 2026-07-21
- FutureCaribbean’s Buildathon offers $50K, H200 compute, and an NYSE pitch — HeyAmit_ · 2026-07-21
- A new series tests which data-science workflows can run on GPUs today — pandeyparul · 2026-07-21
- Former AWS operator says Bedrock margins can beat SageMaker as agentic AI lifts CPU demand — RihardJarc · 2026-07-21