Hugging Face datasets gets major streaming shuffle speedup via optimized Arrow C++
lhoestq · x · 2026-09-07
Quentin Lhoest, author of Hugging Face's datasets library, credits a Google Cloud Storage team contribution for a big speedup when streaming datasets with approximate shuffling. Arrow batches are now buffered, concatenated, and row-shuffled entirely using optimized Arrow C++ operations instead of slower Python-side handling.
Related event: Hugging Face datasets gets major streaming shuffle speedup(2 posts)→
More from Infra
- Tech explores Argentina's Patagonia for mega data centers — pstAsiatech · 2026-09-07
- Average Korean wedding costs as much as an NVIDIA GB300 DGX Station, and people are rethinking marriage — alvelda · 2026-09-07
- Agents on a 16GB MacBook Air M5: Qwen 3.5-9B 4-bit can't yet build a working app — Fluid-Author-9566 · 2026-09-07
- Gary Marcus questions Astra's $1B training bill as scaling data goes dark — GaryMarcus · 2026-09-07
- You're GPU rich when you count in nodes, not GPUs — capetorch · 2026-09-07
- Own or rent? A practical guide to open-weights LLMs vs frontier APIs — spilldahill · 2026-09-07