Human data is bottlenecked by expert labelers' taste, and synthetic data may only boost demand
joecole · x · 2026-09-25
The author argues human training data is bottlenecked by elite SPLs — subject-matter experts whose 'taste' for useful vs. slop tokens compounds with experience and commands outsized pricing power, unlocking multiples of lab data spend. Data firms that don't train SPLs stagnate. Even if synthetic data takes over, a Jevon's paradox effect could increase demand for these 'token surveyors'. The market is still early.
More from Infra
- Smartphones eat ~30% of global DRAM and NAND supply — the fix? Stop yearly phone releases — AlpinDale · 2026-09-25
- TileRT and AMD hit 469 tok/s decode on GLM-5.3 with vLLM on 8x MI355X, 40% faster than GB300 — vllm_project · 2026-09-25
- AI now beats humans at some TPU design tasks, but is still seen as just a tool — burny_tech · 2026-09-25
- More Budget 4-GPU Inference Tricks: x8 Splitters and m.2-to-x4 Adapters — TheZachMueller · 2026-09-25
- Running a 27B model locally on 2x RTX 5090 with vLLM — piddlefaffle12 · 2026-09-25
- CLion 2026.2.3 adds NVIDIA CUDA Tile C++ support with dedicated inspections — blelbach · 2026-09-25