Task Scheduling as a Bandit Problem: FLD Paper Details Continuous-Time Bandit in Human Motion Learning
breadli428 · x · 2026-09-30
Responding to a question about whether the approach resembles a Boltzmann-based bandit algorithm, the author confirms that the task scheduler (teacher) making decisions from learning-progress signals is indeed a bandit problem.
An extension to the continuous-time bandit in human motion learning is exemplified in Section A.2.6 of the FLD paper.
More from Research
- LlamaIndex's Jerry Liu and Snorkel AI on why evals and RL environments remain unsolved — ajratner · 2026-09-30
- Phylo Partners with Michael J. Fox Foundation on AI Agent Target Atlas for Parkinson's — KexinHuang5 · 2026-09-30
- Claude Opus 4.5 beats human at custom wrap-around crazyhouse chess variant — Defiant_Ranger607 · 2026-09-30
- Tempo (UIST 2026): A computer-use agent that reasons about your long-term goals — kenziyuliu · 2026-09-30
- Target Atlas: Hundreds of Biomni Agents Synthesize Drug Target Evidence in Parallel — KexinHuang5 · 2026-09-30
- JevBench Author Plans to Benchmark OpenAI's New Decisions API — airesearch12 · 2026-09-30