The Bayes Bandit: A Mathematical Take on Curiosity in Reinforcement Learning

CatAstro_Piyush · x · 2026-09-04

Francesco Sacco published an interactive blog post, 'The Bayes Bandit: What is Curiosity? Mathematically?', arguing that a good mathematical framework of curiosity could make dataset curation obsolete. Part one of a planned series, it starts from the K-armed bandit problem — Gaussian rewards with unknown means and variances, limited pulls — and walks through the explore/exploit trade-off with hands-on interactive figures showing how Bayesian modeling lets each arm's distribution emerge from evidence. Code is open-sourced on GitHub; the next installment will scale to a full chess engine.

Original post →

More from Research

Research channel →