Proposal: reward 'time to human grokking' in RL environments, especially math

PeterHndrsn · x · 2026-10-12

PeterHndrsn argues RL environments — math especially — should include 'time to human grokking' as part of the reward signal: rewarding only correctness pushes models toward solutions humans can't follow, while factoring in human comprehension cost would yield more readable, transferable solutions. This requires better understanding how humans actually learn.

Original post →

More from Research

Research channel →