Classic revisit: Alex Graves' 2016 Adaptive Computation Time paper pioneered letting networks ponder

peterjliu · x · 2026-09-03

Sharing the classic paper on adaptive compute — Alex Graves' 2016 "Adaptive Computation Time for Recurrent Neural Networks." ACT lets RNNs learn how many computational steps to take between input and output, with minimal architecture change, deterministically and differentiably. It dramatically improved performance on parity, logic, addition and sorting tasks, and while gains were modest on character-level language modeling, computation demonstrably concentrated on hard-to-predict transitions — an early precursor of today's "thinking" and adaptive-compute reasoning models.

Original post →

More from Research

Research channel →