Hugging Face Datasets 5.1.0 fixes streaming bugs for multi-worker PyTorch DataLoaders
lmoroney · x · 2026-10-06
Hugging Face shipped Datasets 5.1.0 with major fixes for iterable (streaming) datasets: take/skip counts now hold when DataLoader workers outnumber shards, epoch iteration works after skipping/taking a subset, shuffling preserves source shards, and distributed skipping shards sources correctly. New splitdatasetbynode strategy argument; adds Vortex columnar format plus FASTA, FASTQ, GenBank, PDB, mmCIF for genomics and molecular structures. Users streaming training data are advised to upgrade and verify per-epoch counts.
Related event: Hugging Face Datasets 5.1 Adds Vortex and Biomolecular Formats(3 posts)→
More from Infra
- GPUs idle waiting on memory while agent task lengths double every 4 months — alex_verem · 2026-10-07
- GPUs idle 85-95% waiting on memory: why the whole AI stack is being rebuilt — alex_verem · 2026-10-07
- AWS Breaks Down ISO/IEC 42005:2025 AI Impact Assessments — AWS ML Blog · 2026-10-06
- AWS Details Four-Layer Governance Model for Shared SageMaker HyperPod Clusters — AWS ML Blog · 2026-10-06
- Marvell moves into custom silicon management as modern BMCs get overloaded — BenBajarin · 2026-10-06
- Amazon Adds Studio UI Management for SageMaker HyperPod Spaces — AWS ML Blog · 2026-10-06