Essay
The Moment You Measure It
We want a world that honors contribution, and the first serious thing anyone says back to us is the most damaging. The moment you measure it, you ruin it. Name the thing you want more of, attach stakes to the number, and people stop producing the thing and start producing the number. It happens every time. It has a law named after it. And if it is true without exception, then the whole project — make impact, not money, the measure of a life — is dead on the table, because a measure of a life is exactly the kind of measure this law destroys.
We did not get this objection from a critic. It is the one that has cost us the most sleep. So we will give it to you at full strength, walk through the cases where it has flattened well-meaning systems, and only then tell you what we think it actually proves. The short version: it does not prove “don’t measure.” It proves something sharper — that the unit of impact must be emergent, the way money is, never a published score that a committee keeps. If we have to measure it, it is already not usable. That is not a dodge. It is the hardest design constraint we have, and the one we are least willing to relax.
The law, stated honestly
In 1975 the economist Charles Goodhart, writing about why the Bank of England’s monetary targets kept misbehaving, put it like this: “Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes” (Goodhart 1975, via the standard account). A relationship that held quietly while no one was watching breaks the instant you lean on it. The anthropologist Marilyn Strathern, studying audit culture in British universities, gave it the form everyone now quotes: “When a measure becomes a target, it ceases to be a good measure” (Strathern 1997). Her own gloss is worth keeping: once a 2.1 degree became the thing to aim for, she wrote, it became “the poorer … as a discriminator of individual performances.” The target ate the signal.
It is not one law but a small family, and they converge. The psychologist Donald Campbell stated the social-science version in 1979: “The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor” (Campbell 1979, via the standard account). And from macroeconomics there is the cousin, the Lucas critique: the statistical regularities a policymaker relies on are themselves built out of people’s expectations, so the act of exploiting them changes the people and dissolves the regularity (Lucas 1976). Three fields, one finding. A measure is a description of behavior that was produced without the measure in view. Put the measure in view, make it pay, and you are no longer measuring that behavior. You are measuring a new behavior — the one aimed at the measure.
We are not going to argue with this. We think it is one of the truest things social science knows. Our disagreement, if we have one, is only about what follows from it.
Where it has actually done damage
Abstract laws are easy to wave away, so look at the bodies.
Start with schools, because the impact case and the testing case rhyme. When test scores become the thing that decides a teacher’s standing or a school’s survival, some fraction of people stop teaching and start manipulating the score. Brian Jacob and Steven Levitt built an algorithm to catch the sharpest version — outright answer-tampering — in Chicago’s public schools, and estimated that serious teacher or administrator cheating showed up in at least 4 to 5 percent of elementary classrooms in a given year, with the rate rising “strongly” in response to even small increases in the stakes attached to the test (Jacob & Levitt 2003). That is the floor — only the cheating crude enough to leave a statistical fingerprint. It does not count teaching to the test, narrowing the curriculum, or quietly steering weak students away from test day, all of which are responses to the same pressure and none of which trip the detector.
Then medicine, where the gaming is more disturbing because the metric was supposed to protect people. England spent the 2000s running its health service on targets backed, as the political scientists Gwyn Bevan and Christopher Hood put it, by “an element of terror” — and they documented what the targets produced alongside the genuine gains: ambulance crews recorded as arriving inside the eight-minute window they had missed, patients held in queuing vehicles or on trolleys so the official A&E clock would not start, waiting lists massaged at the boundary that was being watched (Bevan & Hood 2006). The number improved. Some of what the number was standing in for did not. Worse still are the surgical scorecards. When New York and Pennsylvania began publishing risk-adjusted cardiac-surgery mortality, the economists David Dranove, Daniel Kessler, Mark McClellan and Mark Satterthwaite found that surgeons and hospitals responded partly by avoiding the sickest patients — the ones most likely to die and so to spoil a record. Their verdict is blunt: the report cards “led to higher levels of resource use and to worse health outcomes, particularly for sicker patients,” and “on net … decreased patient and social welfare” (Dranove et al. 2003). A metric built to make care safer made it, for the people who needed it most, more dangerous.
And finance, where the gaming was industrial. Wells Fargo set its branch staff aggressive cross-selling targets — products per customer — and treated the number as the thing to maximize. Employees hit it by opening accounts customers never asked for: roughly 1.5 million unauthorized deposit accounts and around 565,000 credit-card applications, by the regulator’s account, which is why the bank paid a then-record $100 million penalty to the Consumer Financial Protection Bureau in 2016 (CFPB 2016). “Accounts opened” was meant to track customers served. Pressed hard enough as a target, it tracked fraud.
Science is not spared, and this one should worry us most, because science is the closest thing we have to a working impact economy — a community that runs on conferred esteem rather than money. Once the h-index and citation counts became the numbers that decide hiring and tenure, the rest followed: self-citation padding, “citation cartels” in which clusters of authors cite each other far beyond what the work warrants, coercive citation by editors. The literature names the mechanism for what it is — Goodhart’s law applied to itself: a metric meant to track influence, once it becomes the target, starts tracking the manipulation of influence (Ioannidis-adjacent review of citation gaming, Academic Questions 2021). If even the esteem economy we most admire corrodes when it lets a score become the prize, we cannot pretend our own would be immune.
One more case, and we have to handle it carefully, because it is the most famous and the least documented. The Soviet nail factory: told to produce nails by weight, the plant makes one enormous useless nail; told to produce them by count, it makes a heap of tiny useless ones. It is the perfect illustration, and we owe you the truth about it — it is a parable. The vivid version traces to a satirical cartoon in the Soviet magazine Krokodil, and scholars who have gone looking for the original cannot reliably find it. What is documented is the duller real thing behind it: the economist Alec Nove recorded that Soviet sheet glass, planned in tons, came out too heavy, and when the plan was switched to square metres it came out too thin — the same distortion, minus the cartoon (Nove, The Soviet Economic System, 1977, via the standard account). We tell you which is the joke and which is the evidence, because a movement that quoted the joke as fact would be doing the exact thing — optimizing a vivid number over the truth — that it claims to oppose.
So: schools, hospitals, banks, science, the planned economy. The objection is not a worry. It is a track record. Anyone who proposes to honor impact and waves this away has not understood what they are up against.
The thing the cases have in common
Now the turn — and it is not a softening, it is a diagnosis. Look back at every case and ask what was actually present at the scene of the damage. In each one there was a specific published number, declared in advance to be the thing that pays. Test scores decide the teacher’s job. The eight-minute clock decides the trust’s rating. The mortality scorecard decides the surgeon’s reputation. Products-per-customer decides the bonus. The h-index decides the chair. The number was named, the number was central, and the number was the target. People did precisely what the system told them to do; they optimized the declared metric. The gaming was not a betrayal of the design. It was the design, working as specified.
That is the cell where the corruption lives. Not “measurement” in some cosmic sense — a single, central, declared metric that a controlling authority defines and rewards. The instant such a number exists, three things follow as night follows day. It can be reverse-engineered, because it is published. It is worth gaming, because it pays. And it collapses the rich thing it stood for into the one dimension it can count, because that is all a metric can do. Goodhart, Campbell, and Lucas are, read carefully, not warnings against knowing how the world is going. They are warnings against a particular institutional object: the central target.
Which raises the obvious, dangerous question. Money is a measure of value. Money is chased ferociously. Why has money not collapsed under its own Goodhart’s law? Why is “net worth” still, after centuries of the most intense optimization pressure any number has ever faced, a roughly meaningful signal of command over resources?
Here is the answer, and it is the whole essay. Money works as a yardstick because no one sets it. There is no committee that decides your net worth, no published formula you can satisfy, no central scorekeeper to capture. Your wealth is the emergent residue of millions of uncoordinated transactions, each one a tiny local judgment by someone deciding what your work was worth to them. The price of a thing, as Hayek argued in 1945, is a piece of distributed knowledge that no single mind possesses and no planner could compute — it emerges from everyone acting on what only they know (Hayek 1945). Money was never designed as a metric and handed down; it arose, as Carl Menger showed, out of countless people independently reaching for a more tradeable good, until one emerged as the common medium without anyone choosing it (Menger, Principles of Economics, 1871, via the standard account). That is exactly why you cannot game it in the Goodhart sense. There is no central definition to satisfy. To raise your net worth you have to actually persuade many independent people to part with real resources for what you offer — and “persuade many independent people that what you did had value” is not a corruption of the goal. It is the goal. The only way to win the number is to do the thing the number is for.
Money has its own catastrophes, and we are not here to praise it — that what it measures is so thin is the entire reason this movement exists. But its mechanism, the emergent and decentralized part, is the property we have to steal. The disease is the central target. The cure is not “no measurement.” The cure is a unit that is conferred the way a price is conferred — from the edges, by many, with no one in the middle holding the definition.
The rule the law hands us
So the objection, taken at full force, does not bury the project. It writes its strictest law, and the law is a refusal:
The unit of impact must be emergent and peer-conferred, and there must be no central published score. No authority defines it. No formula yields it. No checklist certifies it. The moment a committee publishes “contribution = these five things, weighted thus,” that formula is a target, the target will be gamed, and we will have rebuilt — with worse tools and grander pretensions — the exact machine that failed in every case above. A checklist is a target. We will not publish one. This is the operational meaning of if we have to measure it, it’s already not usable: not that recognition is unknowable, but that the instant it takes the form of an official number to be hit, it has stopped being recognition and become a thing to defeat.
Concretely, the unit has to inherit money’s defenses and add the ones money lacks. Peer-conferred — it comes from the many who witnessed the work, never from a central scorer, so there is no definition to reverse-engineer. Plural — many communities recognizing many kinds of contribution in their own idioms, never a single global number, because a single number is the one most worth capturing and the one that flattens hardest. Decaying — it fades unless renewed by fresh contribution, so it cannot be banked into a permanent position that then defends itself, the way a citation count or a fortune ossifies. And non-convertible — conferred, carried, never cashed; the instant standing converts cleanly and reliably into spendable advantage, it starts behaving like a price, and the whole crowding-out literature we examined in the previous essay comes back to corrode it. Conferred, carried, never cashed. Each of those four is a wall against a specific failure in the cases above, not a slogan.
The hard part, which we will not hide
We have to say plainly where this leaves us exposed, because the honest objection has a second half and it is sharp.
If you refuse a central metric, how do you recognize contribution at all? You have rejected the only mechanism — a defined score — that obviously scales, audits, and explains itself. What is left looks alarmingly soft. And it is soft. Our answer is distributed peer conferral: recognition that emerges from many local judgments by people close enough to the work to see it, aggregated by no one, owned by no one, the way a reputation in a craft or a standing in a scientific field has always formed — slowly, plurally, from below. It is closer to how a neighborhood knows who its quietly indispensable people are than to how a leaderboard ranks them.
We will not pretend this is clean. Distributed conferral has its own failures, and we owe them to you as squarely as we owed you the gaming cases. It can run on popularity instead of substance — the loud rewarded over the load-bearing. It can entrench in-groups, who confer on each other. It is slower, noisier, and far harder to audit than a number; you cannot put it on a dashboard, which is precisely why it resists capture and precisely why it is hard to govern. We think these are the right failure modes to be fighting — failures of an open, plural, human process rather than the cold, total failure of a captured central metric — and we think design can blunt them: weighting conferral by the cost and credibility of the giver so clout is not free, bounding it to real persons, letting it decay so no clique’s verdict calcifies. But “this is the better class of problem and here is how we’d attenuate it” is not “this is solved.” It is not solved. We are claiming the failure modes of emergence are survivable and the failure mode of a central target is not — and that claim has to be earned in practice, not asserted in an essay.
So we will tell you the boundary of what anyone has shown. No one has run a society-wide, emergent, peer-conferred recognition system as a standing institution and watched it hold for a generation. We are not reporting a result; we are proposing an experiment, with the central-score prohibition welded in from the first day so the most common way these things die is foreclosed before it can start. Proposal, not proof. We would rather you held us to that than believed something cleaner than the truth.
What we are sure of is the negative, and it is enough to steer by. Every documented disaster in this essay shares one ingredient — a central number, declared in advance to be the prize — and the thing that has best resisted that disease for the longest time, money, lacks exactly that ingredient. That is not a coincidence we can ignore. It is the design brief. Build the recognition of impact the way value is actually discovered: from the edges, by the many, with no one in the middle holding the score.
Make impact, not money, the measure of a life. Confer it, never compute it — because the moment you compute it, it is no longer the thing you meant.
Sources
- Goodhart (1975), “Problems of Monetary Management: The U.K. Experience”; original formulation, via the standard account — https://en.wikipedia.org/wiki/Goodhart’s_law
- Strathern (1997), “‘Improving ratings’: audit in the British University system”, European Review 5(3), 305–321 — https://gwern.net/doc/statistics/decision/1997-strathern.pdf
- Campbell (1979), “Assessing the Impact of Planned Social Change”, Evaluation and Program Planning 2(1); formulation via the standard account — https://en.wikipedia.org/wiki/Campbell’s_law
- Lucas (1976), “Econometric Policy Evaluation: A Critique”, in The Phillips Curve and Labor Markets — https://ideas.repec.org/a/eee/crcspp/v1y1976ip19-46.html
- Jacob & Levitt (2003), “Rotten Apples: An Investigation of the Prevalence and Predictors of Teacher Cheating”, QJE 118(3), 843–877 — https://academic.oup.com/qje/article-abstract/118/3/843/1943009
- Bevan & Hood (2006), “What’s Measured Is What Matters: Targets and Gaming in the English Public Health Care System”, Public Administration 84(3), 517–538 — https://researchonline.lse.ac.uk/16211/
- Dranove, Kessler, McClellan & Satterthwaite (2003), “Is More Information Better? The Effects of ‘Report Cards’ on Health Care Providers”, Journal of Political Economy 111(3); NBER Working Paper 8697 — https://www.nber.org/papers/w8697
- Consumer Financial Protection Bureau (2016), “CFPB Fines Wells Fargo $100 Million for the Widespread Illegal Practice of Secretly Opening Unauthorized Accounts” — https://www.consumerfinance.gov/about-us/newsroom/consumer-financial-protection-bureau-fines-wells-fargo-100-million-widespread-illegal-practice-secretly-opening-unauthorized-accounts/
- Teixeira da Silva (2021), “Citations and Gamed Metrics: Academic Integrity Lost”, Academic Questions 34(1) — https://files.eric.ed.gov/fulltext/EJ1333169.pdf
- Nove (1977), The Soviet Economic System (glass-by-tonnage example); the Krokodil nail cartoon is a parable of uncertain provenance — https://www.lesswrong.com/posts/YtvZxRpZjcFNwJecS/the-importance-of-goodhart-s-law
- Hayek (1945), “The Use of Knowledge in Society”, American Economic Review 35(4), 519–530 — https://kysq.org/docs/Hayek_45.pdf
- Menger (1871), Principles of Economics (spontaneous emergence of money), via the standard account — https://www.minneapolisfed.org/article/1992/hayeks-legacy-of-the-spontaneous-order