How data lies

Risk Models AI

Does data lie to us when it captures the measurable but not the real? Yes, and it tells three lies. It omits, projects and flatters, misdirecting and leading us astray.

We are becoming ever more data-driven. We invest using models, regulation is evidence-based, and supervision runs on reported numbers, to mention just a few. All fantastic, as data-driven decisions are objective, even scientific, and objectivity and science are hard to argue with.

Some go further. Train AI on enough data, the argument runs, and it forms a world model, an internal representation of everything that matters, along with all the causalities. Nothing important is ignored when AI is brought to bear.

Except it isn't like that. The numbers can be accurate, pass every check and carry the full authority of measurement, and still deceive us. The reason is that data captures the measurable, and the measurable is often, perhaps most often, not what matters most. This is what the sociologist Daniel Yankelovich called the McNamara fallacy, the most extreme form of which is when we dismiss what can't be measured as unimportant.

Data tells three lies.

1. The lie of omission

The picture below shows the distribution of aggregate economic outcomes, from the terrible to the excellent. Almost all our data comes from the middle, the day-to-day fluctuations. The right decisions and the mistakes come from the tails, where the crises and the long-run outcomes live, and are more often than not missing from the data. And when it comes to crises, since every crisis is unique in detail, we make mistakes when using past crises as a guide for reacting to future ones, just as the financial authorities did in 2020, having drawn the wrong lessons from 2008.

Data describes the middle of the distribution in great detail and presents that as a description of the world. What it leaves out is what we care about. The lie of omission.

2. The lie of projection

Models built on data from the middle cope with the scarcity in the tails by making numbers up. Ask a riskometer for the likelihood of a once-in-a-decade crash and it obliges. Plug in a probability and out comes an answer, for once-a-decade events, once-a-century events, even once-a-millennium events. Estimating such numbers honestly would take centuries of data we do not have.

The model was never trained on anything like the events we ask it about, so it invents them. The output looks scientific, and the gap in the tails is papered over with confident numbers.

Take the league table of GDP per capita in 1926 from the Maddison Project Database (Bolt and van Zanden 2020). With rank to the right.

Argentina ranked 11th, richer than Sweden and just behind Germany. Korea ranked 50th, Taiwan 45th and Malaysia 36th, near the bottom. An analyst in 1926, armed with a long span of high-quality data, would likely have predicted continued Argentine prosperity and Korean poverty.

The data was accurate. What it could not contain were the wars, the coups, the institutions and the political choices that drove the reversal, because those had not happened. The outcomes that matter are made by future humans. Good luck predicting those.

Nobody knew how the policymakers of the day would respond to the depression, the wars and the political pressures that followed, and neither did they. The reaction functions of the key decision makers are unknown, unknown-unknowns even, and they determine the outcomes that matter.

Take Argentina. In 1926, Juan Perón was an unknown army captain. Two decades later he was the president of Argentina, and his policies of protection, nationalisation and redistribution would lead it to ruin. No 1926 dataset contained him, and no model estimated on Argentina's golden decades could have anticipated the choices he would make.

Technically, what is missing is the reaction functions. That is why AI world models fall short. However much data the models absorb, the reaction functions are not in it, because the decisions have not been made yet and the people who will make them are unknown.

This is why the world models miss the most important causalities. They will be models of the world as it has been, not the world as it will be.

3. The lie of flattery

Data tells us that what we are measuring is what matters. It validates us. This is the McNamara fallacy again, taken to the extreme, letting us believe that what can't be easily measured really doesn't exist. I know almost nobody would say that out loud, but plenty act as if they do.

When we regulate and risk manage by models, it is, in part, because measured risk is objective and defensible while judgement is not. Bank capital, stress tests, risk management and supervision then all rest on numbers that flatter the powers that be.

Take Dexia. In July 2011, it passed the European Banking Authority's stress test with a comfortable margin under their adverse scenario, more than double the pass mark and among the strongest of the 90 banks tested. Three months later it collapsed and had to be rescued by the Belgian, French and Luxembourg governments. The numbers flattered Dexia, its supervisors and the stress testers alike, right up until the moment they didn't.

Who is doing the lying?

Omission, projection, flattery. Three lies that point our models, our stress tests and our regulations at the best-measured and least significant part of the decision space, while the critical factors go unwatched. When the significant event arrives, we are surprised.

But then collective failure covers individual failure. "I am not to blame, it was the system, not me."

One problem is that a measure that drives regulation becomes a target. Goodhart's law then applies. A measure that becomes a target ceases to be a good measure. Banks can manipulate the forecasts, picking the assets and the riskometers that deliver the returns they want at low measured risk. Measured risk improves while actual risk deteriorates.

Strictly speaking, though, data tells no lies. It is just numbers in a database. We are the ones who read the middle as the whole, project the sample onto the future and equate the measured with the important.

Data does not lie. We lie to ourselves with data.