Strategies for Hyperparameter Tuning on Massive Datasets
Beautiful-Expert-156 · reddit · 2026-07-10
A developer hit a hyperparameter tuning bottleneck while classifying cell types using a dataset of 4.3 million cells and 512 features. Due to the massive training set, each tuning iteration takes an excruciatingly long time, even with an H100 GPU. They tried Optuna and implemented subsampling by extracting only 15% of the training data for tuning, but remain unsure of its robustness. Consequently, they are seeking community advice and solutions for tuning on large-scale datasets.
More from Research
- Draft paper uses Markov-chain eigenfunctions to build partitions and speed up sampling — michaelchchoi · 2026-07-21
- Autoresearch proposes packaging ML runs as studies with questions, analysis, and code diffs — morgymcg · 2026-07-21
- GitHub repo adds lightweight ternary QAT for Prism-ML Bonsai models — terminoid_ · 2026-07-21
- Qdrant co-hosts a Munich meetup on search, retrieval, and agentic RAG on July 23 — qdrant_engine · 2026-07-21
- GigaChat Audio targets long-form audio grounding with timestamps across 120-minute inputs — ai-sage · 2026-07-21
- Paper models Transformer components as stochastic geometry and tests five architectures — Zhihua Liang · 2026-07-21