1,400 people, 14,000 CAPTCHAs: click reCAPTCHA at 3.7s, signup context adds 57%

An Empirical Study & Evaluation of Modern CAPTCHAs

Andrew Searles, Yoshimichi Nakatsuka, Ercan Ozturk, Andrew Paverd, Gene Tsudik, Ai Enkoji

cs.CR

2023-07-22

A 1,400-person study of unmodified live CAPTCHAs finds click reCAPTCHA median 3.7s, contextualized solving up to 57.5% slower, and 30% abandonment.

What problem this solves

CAPTCHAs have been in an arms race with machine learning for two decades. Distorted text was already solved at over 99% accuracy by automated tools in 2014, so sites moved to image selection, sliders, mini-games, and behavioral scoring. Older user studies mostly tested newly proposed schemes or reimplemented text puzzles, not the unmodified widgets people hit on live sites. Researchers at UC Irvine, ETH Zurich, Microsoft, and LLNL asked a plainer question: how long do current deployed CAPTCHAs take, how do people rate them, and how often do they walk away.

Method

They manually inspected the Alexa Top 200. 185 sites had some account-creation flow; 142 could be completed. Each site was visited in Chrome Incognito and in Tor. reCAPTCHA appeared on 68 sites (34%), sliders and distorted text on 14 each (7%), game-like puzzles on 9 (4.5%), invisible "captcha" strings in page source on 12 (6%), and hCAPTCHA on one site. BuiltWith's scan of 673 million sites puts reCAPTCHA at 97.3% of CAPTCHA deployments and hCAPTCHA at 1.4%, so the later user study covers the shapes almost everyone actually sees.

The main study recruited 1,000 Mechanical Turk workers. Each solved ten unmodified live CAPTCHAs in random order: reCAPTCHA v2 at "easiest for users" and "most secure", Arkose rotation and upright-selection, easy and hard hCAPTCHA, a Geetest slider, and three distorted-text styles. Half were told they were in a CAPTCHA study (direct). The other half thought they were creating accounts; the CAPTCHA appeared after submit (contextualized). A second study of 400 people measured abandonment. Across both studies, 1,400 people produced about 14,000 solves. Pay was later doubled to roughly $7.80-$8.20 per hour. Solving time is display-to-submit.

Results

Click reCAPTCHA had the lowest median at 3.7 seconds, with little gap between easy and hard settings. Distorted text came next. Games and sliders had higher medians and wider spread. Mean solving time across types ran from 3.6 to 42.7 seconds.

Mean preference scores sat between 2.76 and 3.94. Click was fastest and best liked, but games and sliders scored well despite taking longer. hCAPTCHA ranked last overall, yet more than 31% still gave it a 4 or 5; 18.9% gave the favorite click widget a 1 or 2. Time and liking do not line up.

Context changes the clock. Direct was faster for every type but one. Easy click reCAPTCHA rose from 3.1 to 4.9 seconds (57.5%). Arkose rotation added 10 seconds (56.1%). The average lift across types was 26.7%. Hard hCAPTCHA, already the slowest median, did not move. Age added 0.09 seconds per year on average, steeper than the 0.03 seconds per year Bursztein et al. reported for text in 2010. Self-reported education did not track speed, which contradicts that earlier "PhDs are faster" result.

TypeHumans (sec / accuracy)Bots in the literature
reCAPTCHA click3.1-4.9 / 71-85%1.4s / 100%
Distorted text9-15.3 / 50-84%under 1s / 99.8%
reCAPTCHA image15-26 / 81%17.5s / 85%
hCAPTCHA18-32 / 71-81%14.9s / 98%
Geetest28-30 / n/a5.3s / 96%

In the abandonment study, 174 of 574 starters quit (30%). Of those who quit, about 25% in the direct setting left before the first CAPTCHA; nearly 50% did in the contextualized setting. Contextualized participants were 120% more likely to abandon. Doubling pay cut abandonment in half in the signup setting and sped solves by 21.5%; in the direct setting, higher pay made people slower (+27.4%).

Why it matters

A CAPTCHA on signup or checkout is a conversion tax. 30% of paid workers still left mid-task, and the signup framing made that worse. Click is almost free; image grids and sliders cost seconds to half a minute. On the security side the picture is already awkward: published solvers beat humans on time and accuracy for most of these types. CAPTCHAs still trip some scripts. They do not stop farms or dedicated models.

For people who build bot checks or run usability tests, the method warning is the part that travels. Asking subjects to "solve CAPTCHAs" produces optimistic times, by as much as half. A study that only uses the direct setting should not be quoted as a live-site number.

Limitations

The authors list several. Direct and contextualized also changed workload, pay, attention, and perceived purpose, so the 57.5% gap cannot be pinned on context alone. There was no consent quiz on MTurk. The main study did not log how many people started and left; the 30% figure comes from the follow-up. Unmodified third-party widgets hide per-challenge accuracy, so Geetest and Arkose have no accuracy column. 163 people typed preference scores outside 1-5 and were dropped. Alexa Top 200 is a 2023 snapshot, and only the signup path was exercised, so deployment counts are a lower bound. The sample is "people willing to finish CAPTCHAs", not web users in general. Bot numbers are cited from prior papers; this study did not run its own attacks.

Terms

Source

What people are saying

Related papers

All paper explainers