2026-08-30
Manheim maps Goodhart failures to four causes and a checklist: diversify, hide, specify post hoc, cap maximization, then run a six-step design process.
Metrics, KPIs, and "objective" scores now sit in hiring, medicine, public health, and model training. Campbell and Goodhart already described the failure: once a measure is used for control, the statistical link to the real goal buckles, and people reshape the system around the number.
People who chase the metric often do not come out ahead either. GPA and completion rates track long-term educational goals poorly. Early COVID dashboards elevated whatever was easy to count. Thomas and Uminsky call metric-reliance a fundamental challenge for AI. This Patterns perspective asks a narrower question: what a metric designer can actually do.
Manheim groups failure causes, then lists design habits, metric features, and a mapping table.
Four causes:
Five design habits: demand coherence; think causally, with a theory of change; use structured compromise when values conflict, rather than forcing one score; run a pre-mortem on how the number will be gamed; schedule reviews that watch behavior drift.
Metric features that can be mixed: new data sources, diversification, aggregation, secrecy, post hoc weights, randomization, soft metrics such as peer review, capping maximization (satisficing), and dropping the metric as an incentive when distortion exceeds the gain.
Table 1 scores each move against cost, immediacy, simplicity, fairness, and non-corruptibility. Table 2 is a six-step process: understand the system and its causal structure, write the goals, pick desiderata, brainstorm measures, plan for failure, then lock a review date.
This is a perspective, not a bake-off. The checkable output is the taxonomy, the worked examples, and the two tables.
Negative cases are specific. The H-index is treated as a universal, unbiased comparison of researchers, while "scientific output" is underspecified and Hirsch barely discusses perverse incentives. England's Year 1 Phonics Check was validated as a screen for at-risk readers, then recycled into an accountability score for teachers and schools. Once flight-duty limits exist, airlines pack schedules against the cap; Shorrock's law says a limit on an efficiency measure becomes a target. Tradable emissions credits over-issued, then went slack once the cap was met. DSM criteria written for clinical judgment get reused by insurers as billing codes, which turns diagnosis back into a gameable metric.
Positive cases show that desiderata shift with the setting. Automated vehicle safety wants measures that are valid, feasible, reliable, and non-manipulatable, and it distinguishes leading from lagging indicators. Development impact bonds need outcomes specified up front, resolvable at maturity, and able to hold up under pressure, with clean causal attribution. The Windfall Clause for a first general-AI profit shock is a one-shot trigger that cannot be validated in advance.
Diversification has a blunt arithmetic: if reading and arithmetic are 50% of school measurement each, science, art, and PE are implicitly 0%. Even a crude hours-in-arts measure reduces the pressure to delete those classes. Stacking tests that fail the same way (everyone cheats) does not help, and it steals class time.
For people who run model evals, RLHF, or internal KPIs, this is a checklist, not a new algorithm. A reward model standing in for human preference has the same failure shape as an H-index standing in for scientific contribution: push the proxy hard, and tail correlation dies. Intervene on the wrong node, and the old relationship is gone.
Several moves are immediately usable. Do not promote a diagnostic measure into an accountability metric. Prefer a threshold to unbounded maximization. When diversifying, put the previously unmeasured goals on the score sheet. Run a pre-mortem before the number pays out promotions or cash. If the incentive's net value does not cover the distortion, keep the number as monitoring and do not attach a reward.
This is incremental engineering advice. It does not prove any metric is safe.
Almost no new experiments. The pluses and minuses in Table 1 are the author's judgments, not effect sizes from a controlled study. Secrecy, post hoc specification, and randomization have fairness and feedback costs the paper admits, without a breakeven bound by organization size.
Academic examples dominate. Metrics inside learning algorithms (training loss, benchmarks, reward models) appear mainly in the closing paragraphs, with no mapping onto the reward-hacking experimental literature. The Windfall Clause is still a design sketch.
"Abandon measurement" is stated strongly, but the operational test remains qualitative: readers must judge when distortion outruns gain.