User-authorized web agents are treated as malicious bots; three principles follow

The Agentic Web Requires New Normative Infrastructure

Cameron Pattison, Matthew Boulos, Noam Kolt, Changbai Li, Tiziano Piccardi, Seth Lazar

cs.CY

2026-06-09

User-authorized web agents are treated as malicious bots. The paper proposes delegation, two-way transparency, proportional restriction, plus FTC and legislative paths.

What problem this solves

By late 2025, browser control, authenticated sessions, and cookie handoff had made it realistic for LLM agents to search, compare, buy, and file complaints on a user's behalf. Capability is no longer the main constraint. The constraint is legal and commercial: platforms and CDNs still dump user-authorized agents into the same blocklists as training crawlers and click farms.

The current toolkit is leftover anti-scraper law. robots.txt, user-agent fingerprints, IP reputation, headless-browser tells, "no automated access" clauses in terms of service, the U.S. Computer Fraud and Abuse Act, and the common-law action of trespass to chattels. hiQ v. LinkedIn left logged-out public scraping intact, but confirmed that platforms can ban automation by contract. In Amazon v. Perplexity, a preliminary injunction started treating "authorization" as a conjunction of user consent and platform consent. Developers reply with logged-in Chromium, residential proxies, and CAPTCHA solvers. Both sides spend real resources. Users often cannot even see that their agent is being quietly throttled.

Method

Three interlocking principles try to split a user's delegate from a third-party crawler.

The regulatory sketch is deliberately light. The FTC is a poor vehicle for a general right of agent access: recent rulemakings have been vacated on appeal, and Verizon v. Trinko rejects a general duty to aid rivals. Section 5 of the FTC Act is the narrower opening. Advertising unfettered access while silently degrading agents may be a deceptive or unfair practice that materially affects subscription choices. On the legislative side, the first ask is mutual identification and disclosed restrictions. Broader bans on blanket blocking sit next to AICOA, the ACCESS Act, and the June 2026 discussion draft of Sen. Warner's AI AGENT Act.

Results

This is a position paper. It reports no new experiment. The figures come from cited industry reports and cases.

SignalFigureContext
Agent traffic per user queryabout 20x a human visitplatforms told the authors
Training crawlers vs user-delegated agents, 2025 peakabout 32xCloudflare Radar 2025, Cloudflare-fronted sites only
eBay v. Bidder's Edgecrawler 1.5% of daily traffic, treated as trespass to chattelsN.D. Cal. 2000
Amazon v. Perplexitypreliminary injunction; authorization trending toward user-plus-platformN.D. Cal. 2025
TomWikiAssistClaude-based Wikipedia editor, March 2026; self-identified, then bannedvolunteer review overloaded

Ad-supported sites have a real problem: if agents fetch HTML over HTTP, display ads may never render. The paper's reply is that ad blockers have long been lawful, so platforms cannot freeze one UI monetization model by banning agents. Licensing (News Corp and OpenAI publicly valued above $250 million), pay-per-crawl, and marketplaces such as TollBit and RSL are already in motion. Proportionality cuts both ways: no anti-competitive purge of bots, but cost-based limits remain available.

Why it matters

For anyone shipping a web agent, browser-use stack, or shopping assistant, the live risk is no longer "will the model click the button." It is "will the click be treated as unauthorized access under ToS and the CFAA." If Amazon v. Perplexity's dual-authorization reading hardens, third-party agents live only on a platform whitelist, and user-side interoperability collapses into in-house assistants.

The three principles are a vocabulary for talking to counsel, policy, and CDNs: prove delegated scope, demand published block rules, and insist that limits map to concrete harm. Cloudflare already says not all bots are bad bots. This is the user-centered version of that line. It does not solve alignment, and it does not price the agentic web. It does drag "should a user's agent be allowed online" out of a private arms race.

The practical ask is concrete. Identify the agent. Keep training crawls off the user-task path. Stop treating silent circumvention as a product strategy. If a platform redesigns an interface only to lock out rival agents, with no user benefit, the paper analogizes C.R. Bard's biopsy-gun tweak that rejected competing needles and Keurig's machines that rejected third-party pods: exclusion that makes no economic sense except to exclude.

Limitations

The authors confine the legal map to the United States, because the major platforms sit there and because U.S. rules travel. GDPR, the DMA, and the Data Act appear only in a footnote.

The three principles are anchors from the consumer side, not a finished balancing exercise. The 20x traffic claim is what platforms told the authors. The 32x figure covers only Cloudflare-fronted sites. Neither is a measurement in this paper. The Section 5 theory depends on agent access already being material to subscription choice, which in 2026 is an inference, not a holding. AICOA and the ACCESS Act are named as vehicles; both have struggled to pass.

Enforcement is the harder gap. Training crawlers still dwarf user-delegated traffic. If identity tokens can be spoofed, the "good bot" lane becomes a training pipe. The paper flags an attribution crisis, multi-agent collusion, and capability-skewed negotiation as real, then parks them as neighboring problems without a mechanism. Equating agents with ad blockers also skips a difference: blockers change rendering, agents change whether a human ever arrives and whether an ad is even in the page. How fast alternative money shows up is left without a timeline.

Terms

Source

What people are saying

Related papers

All paper explainers