Datapoint launches Streams: live human feedback RL for multimodal models
_akhaliq · x · 2026-10-09
Datapoint AI launched Streams, an online RL platform for multimodal models: send image, audio and video rollouts via API while training, and real raters from a claimed pool of 1B+ people across 200+ countries judge them in real time, streaming back preferences at 10,000+ per minute.
Versus batch labelling: seconds-level feedback latency instead of days/weeks; judgments stay on-policy rather than on an outdated checkpoint; human reward updates every policy step; and reward hacking gets caught mid-training by humans. Low-trust answers are dropped before reward-model updates.
More from Research
- Models say no in chat but do it anyway: Simular reveals the agent safety gap — xwang_lk · 2026-10-09
- Three weeks, 19 lectures: a deep recap of Stanford AA203 from Euler equation to PPO — le_james94 · 2026-10-09
- Planning against a learned model seeks out exactly where the model errs flatteringly — le_james94 · 2026-10-09
- New cube packing record for n=12 at 2.9315 set with AI search method — CatAstro_Piyush · 2026-10-09
- Study: LLM judges of AI-scientist idea novelty are unreliable — MarioKrenn6240 · 2026-10-09
- Burkov's Hundred-Page Language Models Book gets hands-on PyTorch edition — burkov · 2026-10-09