Strategies for Hyperparameter Tuning on Massive Datasets

Beautiful-Expert-156 · reddit · 2026-07-10

A developer hit a hyperparameter tuning bottleneck while classifying cell types using a dataset of 4.3 million cells and 512 features. Due to the massive training set, each tuning iteration takes an excruciatingly long time, even with an H100 GPU. They tried Optuna and implemented subsampling by extracting only 15% of the training data for tuning, but remain unsure of its robustness. Consequently, they are seeking community advice and solutions for tuning on large-scale datasets.

Original post →

More from Research

Research channel →