TAMER, the first general-purpose RLHF algorithm from 2008, gets a new playable version

PeterStone_TX · x · 2026-09-06

Peter Stone highlights that TAMER, introduced in 2008, was the first general-purpose RLHF algorithm — predating the recent RLHF boom by over a decade. Brad Knox has released a new version that anyone can now experiment with, offering a hands-on look at the original idea of training agents through human feedback.

Related event: TAMER, the First General RLHF Algorithm from 2008, Gets New Open-Source Release(2 posts)→

Original post →

More from Research

Research channel →