
There is a dangerous assumption built into almost every modern institution:
If the numbers improve, the system is improving.
Usually, that is exactly what a measurement system is designed to tell us.
But what happens when producing a good number becomes dramatically cheaper than producing the reality the number was designed to measure?
The answer is uncomfortable.
The metric can improve while its relationship with reality deteriorates.
Two costs determine whether a metric works
Every measured system contains two production costs.
The first is the cost of producing the underlying performance:
doing the research;
delivering the service;
building the product;
making the loan perform;
actually satisfying the customer.
The second is the cost of producing the appearance of that performance:
generating the report;
producing the submission;
obtaining the score;
formatting the evidence;
hitting the reportable threshold.
As long as those costs remain relatively close, measurement works reasonably well.
It is expensive to fake the signal without producing the underlying reality.
But AI is changing that ratio.
Producing convincing evidence is becoming dramatically cheaper.
Producing genuine performance often is not.
When those costs separate far enough, the signal begins to decouple from what it was supposed to measure.
And failure looks like success
This is what makes the Measurement Trap dangerous.
A broken metric does not necessarily deteriorate.
It can improve.
Output rises.
Targets are achieved.
Scores increase.
Dashboards turn green.
From the perspective of management, the system appears healthier.
Because the very difficulty that previously prevented the metric from being optimized has disappeared.
Scientific publishing provides an unusually visible case.
The recorded retraction rate across scholarly literature stood at roughly 0.2% in 2025.
The organization maintaining the retraction database has stated that it believes the correct rate should be around 2%.
The precise estimate is uncertain.
The structural point is more important:
the system responsible for measuring its own errors believes its measurement substantially understates them.
And another signal is even more revealing.
A significant share of recorded failures now involves compromised peer review.
The attack has moved toward the verification mechanism itself.
When the check becomes the target
There is an important difference between bad output and bad verification.
One fraudulent document inside a functioning checking system is a local failure.
A compromised verification layer creates something different:
a system whose report about its own health is produced by the component that failed.
From above, that can look exactly like excellent performance.
A working verification system discovers problems.
A compromised one approves them.
Management sees fewer problems.
Friction declines.
Performance metrics improve.
That is why more data does not automatically solve the problem.
More data passing through the same compromised verification structure can simply produce more confidence in the same distortion.
Our probability assessment
Our base scenario is Metric Churn — 50%.
Institutions respond by replacing measures more frequently.
But optimization becomes cheaper too.
Each new metric therefore has a shorter useful life than the one it replaced.
We assign 25% to a more important structural transition:
The Provenance Turn.
Instead of asking whether an output looks genuine, verification increasingly asks how that output came to exist.
Preregistration.
Workflow records.
Attestation.
Machine-verifiable chains of custody.
Evidence created while the work happens rather than reconstructed afterward.
We assign 20% to Measurement Retreat — institutions publishing fewer statistics that expose their own weaknesses.
And only 5% to verification being restored sufficiently to close the gap.
What should change in your decisions?
For individuals, use a simple test.
Ask:
What did producing this number cost compared with producing the reality it claims to describe?
A rating anyone can generate cheaply is weak evidence.
A self-reported metric is weaker than an independently produced one.
A physically verifiable measurement or audited transaction is different because producing convincing false evidence remains expensive.
For business, identify the metrics that actually drive decisions and grade them by this production-cost ratio.
Then ask a harder question:
Who verifies them?
Who funds that verification?
And can the function being assessed reduce the budget of the function checking it?
If it can, verification is structurally fragile.
For capital, apply the same test to investment inputs.
Audited financials, settled transactions and physically verifiable production belong in a different category from self-reported engagement, pipeline, impact or forward-looking operating metrics.
They may appear beside each other in the same presentation.
They should not receive the same evidentiary weight.
The scarce asset is changing
For the last two decades, the dominant assumption was that information was scarce.
That world is ending.
Information is becoming abundant.
Generation is becoming cheap.
Optimization is becoming cheap.
Imitation is becoming cheap.
The scarce asset therefore moves one layer deeper.
From information to verification.
And eventually from verification to provenance.
The question is no longer simply:
“Is this convincing?”
It is:
“What had to happen for this to exist?”
That may become one of the most important questions for decision quality in the AI era.
Read the full analysis:
THE MEASUREMENT TRAP
What happens to a system that can see everything and cannot tell when it is wrong
THRIVE IN CHAOS
Decision Intelligence for an Uncertain World
Analysis → Forecast → Recommendations
Signal → Meaning → Action → Stability
Forecasts are probability-based analytical assessments, not certainties. This material supports independent judgment and does not constitute financial, legal or investment advice.
