Hugging Face datasets gets major streaming shuffle speedup via optimized Arrow C++

lhoestq · x · 2026-09-07

Quentin Lhoest, author of Hugging Face's datasets library, credits a Google Cloud Storage team contribution for a big speedup when streaming datasets with approximate shuffling. Arrow batches are now buffered, concatenated, and row-shuffled entirely using optimized Arrow C++ operations instead of slower Python-side handling.

Related event: Hugging Face datasets gets major streaming shuffle speedup(2 posts)→

Original post →

More from Infra

Infra channel →