Hugging Face datasets gets major streaming shuffle speedup via optimized Arrow C++

lhoestq · x · 2026-09-07

Quentin Lhoest, author of Hugging Face's datasets library, credits a Google Cloud Storage team contribution for a big speedup when streaming datasets with approximate shuffling. Arrow batches are now buffered, concatenated, and row-shuffled entirely using optimized Arrow C++ operations. He also thanks contributors Li Yonghui, yuxin00j, and Zhixiang Li.

Related event: Hugging Face datasets gets major streaming shuffle speedup(2 posts)→

Original post →

More from Infra

Infra channel →