AI auditing: The Broken Bus on the Road to AI Accountability
Abeba Birhane, Ryan Steed, Victor Ojewale, Briana Vecchione, Inioluwa Deborah Raji
cs.CY
2024-01-26
Survey of 341 academic AI audits (2018-2022) plus journalism and regulators: campus papers mostly measure bias; lawsuits and investigations force recalls, fines and bans.
Everyone now says they do AI audits. The EU Digital Services Act Article 37 asks platforms to hire "independent auditors." New York City Local Law 144 requires yearly audits of hiring tools. Consultancies sell audits as a product. Fairness papers pile up at FAccT. What is missing is a scoreboard: after the evaluation, did anything change? A product redesign, a recall, a fine, a statute?
Birhane, Steed, Ojewale, Vecchione and Raji tighten the definition to four tests. The assessment must be operationally independent of the team that built the system. It must name a concrete target, a real deployment or a widely used open dataset, not a toy model the authors trained themselves. It must measure that target against stated expectations. And it must aim at accountability. Abstract fairness experiments do not count. The paper asks who actually audits, how they do it, and which of those studies produce consequences.
On the academic side they took 2018-2022 proceedings from FAccT, AIES, EAAMO, AAAI, CSCW, IC2S2 and WWW, then keyword-searched the ACM Digital Library and ACL Anthology on six title terms: audit, accountability, case study, bias, fairness, assurance. That yielded N=341 papers: 55 from FAccT, 40 from AIES, 137 from EAAMO/AAAI/IC2S2/WWW, 103 from other ACM venues, 6 from ACL. Science, Nature and PNAS were left out, as were economics and sociology journals.
They tagged four kinds of work: product/model/algorithm audits, data audits, ecosystem audits that look at affected communities and the surrounding sociotechnical environment, and meta-commentary (methods papers plus critiques of auditing itself). Of 337 papers with abstracts, roughly 212 are product-level, 46 data, 15 ecosystem, 64 meta-commentary. A paper that straddles categories is counted once, under the best fit.
Outside academia they sampled public reports and websites from six domains and coded each on motivation, target, harms, institutional position, methods and observable impact. Journalism: ProPublica and The Markup. Civil society: EFF, ACLU, Citizen Lab, Ada Lovelace Institute, Refugee Law Lab, Migration Tech Monitor. Government: the UK ICO and US NIST. Consulting: ORCAA, Eticas, BABL AI. Law: Foxglove, Luminos.Law, AWO. They also note BSR human-rights reviews commissioned by Google and Facebook.
Author keywords already show the academic taste: fairness in 25.1% of keyword lists, bias in 24.7%, accountability in 15.6%. Abstracts make the split sharper.
| Type | Abstracts mentioning bias | Abstracts mentioning accountability |
| Product/model/algorithm (212) | 32.5% | 13.7% |
| Data (46) | 39.1% | 8.7% |
| Ecosystem (15) | 13.3% | 33.3% |
| Meta-commentary (64) | 15.6% | 28.1% |
Six of the 15 ecosystem papers used interviews, workshops or ethnography. Product audits almost never did.
Consequences diverge by domain. ProPublica's COMPAS investigation became the template for algorithmic auditing. The Markup used data donation to probe Facebook ad targeting; the US Department of Justice later made Facebook drop a special-audience tool for housing ads. A liver-transplant matching algorithm that favored wealthy urban patients was scrapped. On the civil-society side, the ACLU won an injunction in K.W. v. Armstrong against algorithmically driven welfare cuts for people with developmental disabilities in Idaho; EFF won a student's civil case against the exam-proctoring vendor Proctorio. The UK ICO fined TikTok £12.7 million for misuse of children's data and Clearview AI £17 million, and ordered Clearview to stop processing UK personal data. Foxglove helped reverse Ofqual's A-level grading algorithm, halt a Home Office visa-streaming algorithm, and force disclosure of NHS-Palantir contracts.
Academic impact is usually invisible at publication. A few cases show up years later: Gender Shades named commercial gender classifiers and those products were later changed; 80 Million Tiny Images was withdrawn; Microsoft's celebrity-face dataset was discontinued after a Financial Times investigation. Consulting work is often opaque. After ORCAA audited HireVue's campus-hiring tools, HireVue said it would drop facial analysis, then drew fire for letting the client set the scope and for spinning the findings. NIST's AI Risk Management Framework entered the US policy conversation; NIST has no enforcement power and has not run an audit with a specific, documented outcome.
The discussion turns that contrast into working rules. Who the auditor stands with, and how power is distributed, predicts impact better than a new fairness metric. Target selection, harm framing and how results are communicated often matter more than the evaluation code. Naming a specific product and a specific demand is more likely to force a response than a diffuse ecosystem scan; ORCAA's HireVue review was criticized as too high-level to support any allegation. Internal auditors have access and get blocked by conflicts and gag rules. External auditors can publish and lack data and legitimacy. Timing (before, during or after deployment) and internal versus external status are not the main drivers.
Some harms do not fit an audit at all: the chilling effect of surveillance, concentration of corporate power. Audits are one tool in a larger accountability kit.
If a compliance checklist says "commission an independent audit from a consultancy," the evidence here is that evaluations paid for and scoped by the target rarely produce observable accountability. The "independent auditors" in DSA Article 37 look, under this paper's definition, more like internal auditors. In the sample, the path that actually changed systems was more often investigative reporting, public-interest litigation, and regulators who can levy fines.
For people still writing audit papers, treat this as a constraint on topic choice. A bias table by itself almost never shows consequences at publication time. Gender Shades worked because it named specific commercial systems and because journalists and regulators picked the findings up.
When a model lab asks a university or a consultancy to serve as a third-party auditor, the institutional brand does not automatically convert into accountability. In this sample, the public impact of those collaborations is often listed as "unclear."
The academic sample is bounded by design: computing-adjacent conferences from 2018-2022, no Nature/Science, no disciplinary social-science venues. Ecosystem audits number only 15. Cross-cutting papers are forced into a single label. The non-academic side is a purposive sample, not a census. Much consulting and law-firm work is marked privileged and confidential, so impact scoring tilts toward already-public success stories.
Impact is coded generously: press coverage counts. Academic consequences often lag, so scoring from documents available at publication systematically undercounts campus work. Keyword counts are not full-text coding; 32.5% of abstracts saying "bias" is not 32.5% of papers that only measure bias.
The title's "Broken Bus" is never explained in the body, so it is not an argument. The paper also never defines a reproducible score for "effective audit"; the cross-domain comparison is qualitative. The appendix points to a public list of the 341 papers and the term-count code; the qualitative labels are not a scoring rubric one can rerun.