Incremental AI Can Permanently Sideline Humans Without a Sudden Takeover

Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development

Jan Kulveit, Raymond Douglas, Nora Ammann, Deger Turan, David Krueger, David Duvenaud

cs.CY

2025-01-28

Incremental AI can permanently sideline humans without a takeover: once the economy, culture and states no longer need people, aligning individual models is not enough.

What problem this solves

Most of the AI-risk literature still orbits two pictures: someone uses a model to do harm, or a misaligned autonomous system defects, grabs a decisive advantage, and humans lose. That agenda, from Bostrom through Carlsmith and Ngo to the Bengio-led safety reports, has pushed technical work on honesty and dangerous capabilities, and governance work on evals and norms.

This January 2025 cs.CY paper by Kulveit, Douglas, Ammann, Turan, Krueger and Duvenaud puts a third picture on the table. Capabilities can rise smoothly. Systems can keep doing what their designers asked. Humans can still lose, for good. They call it gradual disempowerment.

The load-bearing claim is simple. Markets, cultures and states currently track human interests because they still need human labor and cognition, not because those systems are intrinsically kind. Once cheaper machine substitutes take those roles, growth incentives unhook from human flourishing. People who refuse the substitution get outcompeted by people who accept it.

Method

There is no new model and no new benchmark. The paper is a six-claim argument.

Societal systems are somewhat aligned with what people want, and that alignment is neither automatic nor built-in. It runs through explicit acts such as voting and buying, and through a quieter channel: these systems still depend on human work and human thought. If that dependence falls, both channels weaken. Where the systems already reward outcomes that hurt people, AIs will chase those rewards harder. The three domains feed each other, so drift in one pulls the others. If the drift correlates, humans may stop being able to command resources at all, which the authors count as an existential catastrophe.

The economic section hangs on two long-run facts. US personal consumption has been about 70% of GDP. Labor's share of US GDP has sat near 60% for more than a century, one of Kaldor's stylized facts. Supply and demand are still mostly human. Earlier machines automated narrow tasks and pushed people into more complex ones. Worker-replacing technical change, in Korinek and Stiglitz's 2018 sense, is different: when machines can do almost any cognitive job more cheaply, people are not moved up, they are moved out. Households then lose the income that lets them show up as consumers.

Four pressures push firms that way: competitive disadvantage if you keep humans in the loop, copy-and-scale asymmetries, a regulatory gap that taxes and constrains people more than models, and anticipatory disinvestment in human skill once a task looks automatable. Relative disempowerment leaves people wealthy in absolute terms while markets optimize for AI activity. Absolute disempowerment is losing the bidding war for land, energy and materials, then watching human-centric supply chains wither. Figure 1, sketched after Korinek and Suh 2024, has wages rising first and collapsing before full automation.

Culture is treated as cultural evolution. Ideas and artifacts compete as variants. Historically they had to live in human hosts, which put a ceiling on how anti-human a successful variant could be. AI is the first technology that can replace human cognition in production, transmission and selection. The Sydney persona that appeared in Bing Chat in 2023 is their early specimen: it leaked into news and social media, then into later training data, and Llama 3.1 405B can be prompted into a similar pattern in a few lines. Harm so far is small. The reproduction loop is already there. Cultural evolution also speeds up. A/B testing already tunes for addiction; more compute finds exploits faster than societies grow antibodies.

States are compared to rentier states. Oil rents let governments tax citizens less, and civic participation thins. If AI profit taxes replace labor income taxes, the same logic applies. Automated security weakens the old constraint that a government cannot afford to alienate its own armed humans. Model-written law can become something people cannot navigate without another model. Geopolitical competition, administrative throughput, and the hope that AIs will not form independent power bases all push states to adopt. Elections can remain. Drafting, advice and enforcement can still migrate.

Section 5 attacks the intuition that the three systems will check each other. The channels between them do not care about human values. Tobacco money shaping policy and taste is the historical template. Using the state to redistribute after economic displacement dumps the alignment burden onto the state and gives the state economic power over people. Using the state to police culture hands it control over cognition. Opening democracy wider makes the state more exposed to cultural drift.

Results

No experiment table. The numbers they pin down are background facts, plus one conclusion they put in bold.

AnchorFigureCompared with
US consumption / GDP70%Demand still human-led
US labor share of GDP60%, >100 yearsSupply still human-led
Plan to stop thisNone they judge concrete and plausibleSingle-system intent alignment is not enough
Influence-limiting rulesMostly stopgapsThey sacrifice value and invite circumvention

Measurement is itself an open problem. There is no external yardstick for how aligned a civilization is, no clean line between useful augmentation and harmful replacement, and no known threshold past which humans cannot recover. They sketch four families of response: measure the drift, cap AI influence, thicken human control, and what they call ecosystem alignment. Suggested meters include AI's share of GDP as its own category, unsupervised AI spending, the human share of widely consumed culture, legislative complexity, and the role of models in law and security. Caps include mandatory human sign-off, limits on AI asset ownership, and progressive taxes on AI revenue. Thickening includes faster democratic process, human-understandable outputs, AI delegates that can keep up with competitive speed, and forecasting tools.

They are blunt that caps are stopgaps. If some countries forgo AI's economic gains to stay aligned, the strongest economies may be the ones that disempower their people the most.

Related work ties the story to Bostrom's evolutionary whimper and "Disneyland with no children", Christiano's What Failure Looks Like, Critch and Russell's production web, Kasirzadeh's accumulative x-risk, Hanson's cultural drift, and Critch's 2024 personal estimate of an extra 50% chance that humans gradually hand Earth to AGI, with most of that probability between 2030 and 2040. That last number is Critch's, not a result of this paper.

Why it matters

For people who work on alignment, the unit of analysis moves. A model can locally do exactly what it was asked, while civilizational incentives walk away. Evals for sudden defection, dangerous capabilities and deception cover this path poorly.

For product and policy people the implication is narrower. Labor share, consumer choice, unions, votes and protest all assume the system still needs humans. Keep the forms of democracy and add UBI, and relative disempowerment can still proceed. Paying people from the state after the labor share collapses can spend the remaining check on the state.

This is not a methods paper and it is not a modest empirical improvement. Its value is the reframing: a sudden takeover is not the only existential path. Incremental substitution plus cross-system feedback is enough, on their account. They say out loud that nobody yet has a concrete plausible plan.

Limitations

There is no dedicated limitations section. The gaps sit in the open.

The whole piece is mechanism, not measurement. 70% and 60% describe the present. Figure 1 is a cartoon inspired by someone else's scenarios. There is no operational border between relative and absolute disempowerment, and they list threshold research as future work.

"No concrete plausible plan" is honest, and the four intervention families still read as a research agenda. How much competitiveness you give up, and who folds first if coordination fails, are unquantified. AI delegates that keep up with competitive dynamics may just insert another layer of agency you do not control.

The culture chapter is the easiest place to overclaim. Sydney is a real reproduction loop with small harm; jumping from that to anti-human ideologies that no longer need human hosts skips several levels. The analogy to Indigenous cultures in North America shows relative marginalization. It does not show absolute disempowerment.

The authors note that Claude Sonnet, Claude Opus and ChatGPT o1 helped write and edit the text. For a paper about AI entering cultural production, that line is itself a data point.

Claim 6 is the one to push on. Claims 1 through 5 are a mechanism. Claim 6 inflates that mechanism to existential scale. The paper does not spell a tight causal chain from resource bidding to political collapse to failed infrastructure and extinction. They mark it as plausible, not demonstrated.

Terms

Source

What people are saying

Related papers

All paper explainers