First-party social-impact evals average 0.77/3 as third parties fill the gap

Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations

Anka Reuel, Avijit Ghosh, Jenny Chim, Andrew Tran, Yanan Long, Jennifer Mickel, Usman Gohar, Srishti Yadav, Pawan Sasanka Ammanamanchi, Mowafak Allaham, Hossein A. Rahmani, Mubashara Akhtar, Felix Friedrich, Robert Scholz, Michael Alexander Riegler, Jan Batzner, Eliya Habba, Arushi Saxena, Anastassia Kornilova, Kevin Wei, Prajna Soni, Yohan Mathew, Kevin Klyman, Jeba Sania, Subramanyam Sahoo, Olivia Beyer Bruvik, Pouya Sadeghi, Sujata Goswami, Angelina Wang, Yacine Jernite, Zeerak Talat, Stella Biderman, Mykel Kochenderfer, Sanmi Koyejo, Irene Solaiman

ICML)

cs.CY, cs.AI, cs.LG

2025-11-06

186 first-party reports average 0.77/3 on social-impact evals vs 2.64 for third parties; environmental and bias disclosure is falling, and moderation labor is nearly absent.

What problem this solves

Governance already treats evaluations as the main evidence of foundation-model risk and capability. Capability leaderboards are everywhere. Reporting on bias, fairness, privacy, environmental cost, and labeling or moderation labor is not. This ICML 2026 paper pins the question to seven base-level social-impact dimensions from Solaiman et al. and asks, at scale, who reports them, how detailed the reports are, and where the holes sit.

The unit of analysis is the foundation model before it is tied to a specific application. Downstream systems need context-specific rules. What developers can measure, and should disclose first, are properties that do not depend on a deployment.

Method

Three strands run in parallel.

Each artifact is scored 0–3: nothing reported, a vague mention, a number without a usable method, or enough detail to reproduce. Cost categories also credit hardware and resource figures. 390 items were double-annotated; disagreements went to a lead author. A Bayesian hierarchical ordinal model then relates scores to openness, organization type, and year.

Results

First-party reports average 0.77 in detail; third-party evaluations average 2.64. The regression finds the same gap on all seven dimensions.

Sensitive content is the most frequently and most carefully reported category, matching how corporate risk teams think about reputational harm. Data and content-moderation labor is almost missing: 8.06% of first-party reports mention it, and third parties barely compensate. Environmental-cost reporting drops after 2023 Q3; bias reporting follows a similar slide. Interviewees said evaluations usually run only when they help product adoption, move a business metric, or are forced by compliance.

In a stratified provider sample, Google averages 1.30, Meta 1.18, Cohere 1.00, OpenAI 0.74, and Mistral 0.32. Academia leads release-time reporting at 0.77, industry and nonprofits sit at 0.56, government at 0.36. Third-party work clusters on U.S. models, then Chinese ones. Low-visibility models get little scrutiny.

Within-provider histories make the slide visible. Meta reported bias more carefully from OPT through Llama 2 and then faded to vague mentions; Google and Anthropic show a similar thinning. Third parties spend their budget on reproducible behavioral tests. Supply-chain facts that only the developer holds remain the thinnest layer.

Why it matters

Independent labs can fill gaps on bias, harmful content, and performance disparities. Provenance, moderation labor, training cost, and data-center energy still require numbers only the developer holds. Anyone using a model card to decide on adoption is often looking at an empty social-impact column, and that column is getting emptier.

The policy ask is concrete: mandate developer disclosure, give independent evaluators a safe harbor, and build shared infrastructure that aggregates third-party results. The annotation dataset is public.

Limitations

The interview sample is small and thin outside Europe and North America. Provider sampling favors visible models, so the empirical picture may inherit that bias. Scores measure presence and specificity, not whether a reported eval is a good measurement. The seven categories are not exhaustive. The authors treat this as a transparency baseline, not a validity audit.

Terms

Source

What people are saying

Related papers

All paper explainers