The Dynamics of Intelligence Explosions
Toby Ord
cs.AI, econ.TH
2026-08-15
Ord shows RSI hits a finite-time singularity only if generation time shrinks to zero fast enough; otherwise growth stays super-exponential without a vertical wall.
I. J. Good's 1965 intelligence explosion is a loop: a machine that outthinks humans can design a better machine, then another, until human intellect is left behind. The modern label is recursive self-improvement (RSI). Ord uses it inclusively, covering human-heavy R&D, mixed pipelines, and fully automated successor design.
Recent economics-inspired models write capability A as Ȧ = kA^r. When r>1 the solution is a hyperbola that hits a vertical asymptote at a finite time t, a mathematical singularity. Jones's semi-endogenous growth writes Ȧ = δ L^λ A^{1-β}, with λ ≤ 1 for overlapping work and β > 0 for ideas getting harder to find; hold population fixed and even exponential growth dies. Davidson and coauthors restore the possibility by setting AI labour L = C A, compute times efficiency, so the exponent becomes 1+λ-β and can exceed 1.
Ord's claim is that this functional form lumps two different explosions together. Super-exponential growth need not hit a wall in finite time. The neglected control knob is generation time: how long one trip around the feedback loop actually takes.
He drops the power law and studies Ȧ = f(A). Super-exponential growth needs f to be super-linear in A. A singularity needs a global blow-up condition: the integral of 1/f(A) from today to infinity must converge. Convexity is neither necessary nor sufficient.
That split is already visible in Ȧ = A log A. The right-hand side is super-linear, yet A(t) is doubly exponential and the integral diverges. Stacking extra log log factors stays on the same side of the threshold; raising the last factor to a power greater than 1 crosses it. Kurzweil preferred A log A for the same reason: a finite world is a poor source of infinite knowledge.
Feedback loops are not Newton's law of cooling. They take time. In a raw difference equation ΔAn = f(An), no growth rate produces a singularity: finitely many finite increments stay finite. Ord embeds the difference equation in calendar time by giving step n its own generation time Tn.
The singularity theorem splits the problem into two independent gates that cannot compensate for each other. The Zeno condition: Σ Tn converges, so infinitely many loops fit in finite time (1/n is too slow, 1/n^{1.01} is fast enough). The boundlessness condition: Σ ΔAn diverges to +∞, so the per-loop increment cannot shrink too fast.
Gradient is rise over run. In discrete time, shrinking the run is what drives a vertical wall. The rise per step can even stay constant. Doubling time falling toward zero marks super-exponential growth; only generation time marks a singularity. A delay differential equation with fixed lag behaves the same way: fast, never singular.
Solomonoff's toy model makes the point concrete. If AI labour can keep Moore's Law running, a 2-year doubling, then 1 year, then 6 months, produces a singularity in 4 years. Moravec's threshold is that the nth generation must be faster than the first by more than n log n log log n. Eth and Davidson's training-time compression story lands on the same family.
| Model | Super-exponential implies singularity? |
| Ȧ = kA^r | Yes; r>1 is hyperbolic |
| General Ȧ = f(A) | No; double exponentials sit in between |
| Pure difference ΔAn = f(An) | Never |
| Time-embedded difference | Much harder; super-exponential without a singularity is the default |
A sharp pair of examples: Tn = 1/n with tetrational growth still misses a singularity. Tn = 1/n^{1.01} with +1/n per loop hits both gates and blows up.
Choice of scale moves the super-exponential line, not the singularity line, so long as two measures are similar: one unbounded exactly when the other is. Chess progress looks roughly linear in Elo and exponential in Bradley-Terry strength, the log of the same quantity. Mean time between failures, or the METR time horizon, can diverge while other cognitive skills stay finite. Infinite horizon on that benchmark is 100% reliability on its task class, a coordinate singularity, not unbounded intelligence.
Physical floors arrive first. Generation times range from decade-scale EUV-successor loops to second-scale scaffold edits; the short loops cover a thin slice of the pipeline and saturate. If Tn bottoms at T while each loop multiplies A by a constant factor, A(t) settles onto a faster exponential, not a vertical wall. Ord sketches phases: human-speed exponential, then super-exponential while generation time is still falling, then machine-speed exponential, then logistic saturation at some A.
He wants frontier labs to report generation times, especially for pretraining and RLVR post-training. Even a linear speedup to A(10t), a decade of human-only progress each year, already carries most of the danger. The curve need not change shape.
For anyone tracking takeoff, software intelligence explosions, or semi-endogenous RSI models, this is a measurement correction. Monthly growth of 10%, then 20%, then 30% looks hyperbolic inside a power-law model. The same sequence is a double exponential on the order of exp(e^{0.1 t}), not a vertical asymptote. Local elasticity above 1 neither confirms nor rules out an explosion, because blow-up and Zeno are global.
The quantity worth instrumenting is how long one loop takes, not only how much capability rises. Internal evals and regulation that watch elasticity while generation time has a floor will see the model fail long before any intelligence ceiling.
The paper is theory, not a new experiment. It turns "singularities are hard" into two non-substitutable thresholds, and treats super-exponential-but-not-singular growth as the default once discrete loops are modeled honestly.
There is almost no empirics. Ord does not estimate current Tn from training cycles, chip tape-outs, or agent loops, nor how far those times can fall. The appendix table is an analytic catalogue of differential equations, not observations.
A is a placeholder. He notes that intelligence has no agreed cardinal scale, and offers no operational definition. The phase sketch assumes constant-factor gains per loop, with generation time bottoming out before capability does. That is a scenario, not a fit.
The Zeno condition idealizes a countable infinity of cycles. Clocks, causal delay, and fab schedules stop first. Ord already treats t as the point where the model must break, not a forecast of infinite intellect.
The closing A(10t) warning does not quantify which speed is already ungovernable. Risks from RSI outrunning safety work are named in the setup and then left aside.