Google details the first vectorized, performance-portable Quicksort
mococa · hn · 2026-09-17
Google's open source blog describes the first vectorized, performance-portable Quicksort, implemented in its C++ parallel library. The design uses SIMD instructions to parallelize partitioning and comparisons while remaining portable across CPU architectures. The HN thread digs into the engineering details and real-world speedups.
More from Infra
- Crusoe runs 512 AMD MI355X GPUs at 5.75M tok/s in largest MLPerf inference entry — wkmyrhang · 2026-09-17
- Perovskite could lift solar efficiency ceiling from 30% to 45% — and give the US a shot against China — kyliebytes · 2026-09-17
- How mobile and specialization broke homogeneous compute into TPUs, NPUs, and more — blelbach · 2026-09-17
- After the x86 Monoculture: Software Will Suffer for Hardware's Fragmentation Again — blelbach · 2026-09-17
- Hardware veteran: low-precision gains nearly exhausted, true sparsity is AI's next 10x — blelbach · 2026-09-17
- Moore's Law in three eras: from free lunch (1970-2005) to software hell (2015-now) — blelbach · 2026-09-17