Fooled by Best Practice, Part I: The Cemetery You Cannot See
J. McKenney
Paper 1 of five in Fooled by Best Practice. This part makes the argument in ordinary language and assumes no background. Part II sets out the distinction it rests on. Part III gives the method. Part IV gives the engineering response. Part V gives the decision.
Licence: CC BY 4.0. 16 September 2026.
| Field | Value |
|---|---|
| Document ID | WG-02-DT-FBP-1 |
| Slug | fooled-by-best-practice-1-cemetery |
| Working group | WG-02-DT, Digital Twin |
| Series | Fooled by Best Practice, part 1 of 5 published |
| Author | J. McKenney |
| Published | 16 September 2026 |
| Revision | |
| Status | Published |
| Preceded by | |
| Followed by | WG-02-DT-FBP-2, WG-02-DT-FBP-3, WG-02-DT-FBP-4, WG-02-DT-FBP-5 |
Executive Abstract#
In 2001 Nassim Nicholas Taleb argued that the people running financial markets could not tell luck from skill. The industry only ever looked at the traders still standing, and the ones who failed had left, taking their evidence with them.
I have spent more than twenty years inside substations, battery storage sites, rail control rooms and water treatment plants across North America, Australasia and Europe. I read Taleb not as a book about finance but as an accurate description of my own industry.
The sentence I hear most in a critical facility is some version of we have not had a significant incident. It is offered as proof the defences work. It is not. Three years without a serious incident is one route through a vast space of possible routes, and some of the others end with a process out of control, a safety system failing when called on, or equipment destroying itself. That those routes did not happen this time says nothing about how likely they were.
This paper is about that error, the oldest in risk. It names four forms I meet in the field: the architecture diagram that stopped matching the plant years ago; the audit score read as a security posture; the vendor reference list with the failures quietly missing; and the practice nobody questions because it has held for fifteen years. It is not an argument against standards. I use IEC 62443 constantly. It is an argument about what a passed audit is actually evidence of, which is narrower than the industry treats it, and about what must be measured instead.
Abstract#
This paper opens a five-part series applying the epistemology of Taleb's Fooled by Randomness to operational-technology security in critical infrastructure. It argues the sector commits Taleb's survivorship error in structurally identical form: defensive competence inferred from the absence of observed compromise, a single realized history rather than the distribution it was drawn from. Without mathematics, it separates the record of what happened from the space of what could have, gives the arithmetic by which winning streaks arise among identical strategies, and distinguishes two ways evidence goes missing: outcomes never observed, and outcomes observed by parties under no obligation to report them. It names four field patterns, the reference-architecture, compliance, vendor-survivorship and institutional-memory illusions, and argues the quantity being measured is not the one operators believe. The remaining parts give the distinction formally, the method, the engineering consequence, and the allocation decision.
1. Introduction#
1.1 A book about traders that described my industry#
I came to Taleb late and by accident. Fooled by Randomness had been out for years before I read it, and I picked it up expecting a book about markets, which is a subject I have no professional stake in [1].
What I found was a description of my own working life.
Taleb's subject was a room full of fund managers, each of whom believed their returns were the product of method. His argument was that in any large population of managers making essentially arbitrary bets, arithmetic alone guarantees that some will have long winning streaks. Those are the ones who get promoted, profiled and copied. The ones whose strategies failed are gone. They do not appear in the industry's data, they do not speak at its conferences, and their absence is invisible precisely because they are absent.
I had been watching the same thing happen in control rooms for two decades without a name for it.
The language in a plant is different from the language on a trading floor. Nobody in a substation says their Sharpe ratio is strong. They say they passed the audit. They say they are 62443 compliant. They say, most often of all, that they have not had a significant incident. The sentence carries the same weight a track record carries in finance, and it is doing the same work: standing in for evidence that the approach is sound.
The error underneath is identical, and so is the arithmetic.
1.2 What this part does, and what it does not#
This part makes the argument in plain language. There is no mathematics in it, and none is needed. If you can follow the idea that a coin can come up heads five times in a row without being weighted, you have everything required.
The rest of the series goes deeper by design. Part II states the distinction this argument rests on and shows where the map stops matching the territory inside a real plant. Part III gives the method for measuring the thing that is actually at issue, with its mathematics and its limits. Part IV is the engineering case, and it is the most technical of the five. Part V is written for whoever signs the budget.
One thing this series is not. It is not an attack on standards, and I want that clear at the start rather than buried in a qualification at the end. I have spent a large part of my career working inside IEC 62443, and I think the framework encodes hard-won knowledge that would be foolish to discard. The argument here is narrower and it is about inference. A passed audit is evidence of something real. It is not evidence of the thing operators take it to be evidence of.
2. The survivors are the only ones who file reports#
2.1 Ten thousand traders#
Take ten thousand people and give each of them a strategy that is genuinely no better than a coin toss. Each year, half of them do well and half do badly. After one year, five thousand have a winning year. After two, two thousand five hundred have two. After five years, a little over three hundred of them have an unbroken five-year record.
Nothing in that population had any skill. The arithmetic produced the winners on its own.
Now do what the industry does. Interview the three hundred. Ask them what they did. They will tell you, in complete sincerity, about their method, their discipline, their read on the market. They are not lying. They have no way of knowing that their record is what a fair coin looks like when you run it ten thousand times and then only look at the streaks.
Taleb's point was not that these people were frauds. It was that the industry had built an evaluation system that could not distinguish them from people who were genuinely good, and had then used that system to allocate enormous amounts of capital.
2.2 The same arithmetic in a control room#
Now run the same exercise on facilities.
There are many thousands of industrial sites in any given sector. In any given year, only a small fraction of them suffer a serious security incident. That is not because most of them are well defended. It is because a serious incident requires several things to line up at once: an adversary with the capability and the motive, a target selected out of many, a vulnerability in a state where it can actually be reached, and a moment when nobody notices in time.
Most sites, in most years, do not have all of those converge. So most sites, in most years, report no incident.
After three years, the great majority of sites in the sector can say they have had no significant incident. They will say it, in my experience, in a tone that means the question is settled.
It is not settled. It is the five-year winning streak, and there are thousands of them, and the arithmetic produces them whether or not anything in the plant is actually working.
2.3 Two ways evidence goes missing#
There are two separate mechanisms here and they are worth keeping apart, because only one of them is what Taleb was describing.
The first is the one he wrote about. Outcomes that did not occur leave no record anywhere, because there is nothing to record. The plant that would have failed under a slightly different sequence of events did not fail, so there is no incident, no report, and no lesson. In The Black Swan Taleb later gave this its clearest name, silent evidence, and made the point that the cemetery does not publish [2].
The second is particular to this industry and is not a statistical effect at all. A great many incidents in critical infrastructure are observed by someone and then not disclosed. They are handled under non-disclosure agreements, settled quietly with the vendor, contained inside the operator's own incident process, or reported only to a regulator who does not publish. I have been in rooms where an incident was discussed frankly and I have watched the same incident be absent from every public account of that sector for the following year.
The first mechanism means the data is incomplete by nature. The second means it is incomplete by arrangement. Both push in the same direction, and the combined effect is that the public record of what goes wrong in critical infrastructure understates it, by an amount nobody can quantify because the quantity is the missing part.
When somebody shows you a low incident rate for a sector, you are being shown a number whose denominator is honest and whose numerator is not.
3. One sample path#
3.1 What "three years without an incident" establishes#
Here is the claim as carefully as I can put it.
A facility that has operated for three years without a serious security incident has demonstrated one thing: that the particular combination of adversary activity, configuration state, human behavior and timing which occurred during those three years did not produce a realized attack path.
That is a real fact and it is worth having. It rules out the worst case, which is a facility already being exploited.
What it does not establish is the thing it gets used for. It does not tell you how close you came. It does not tell you how many of the alternative routes through those same three years ended differently. It does not tell you whether the defences held because they are sound or because nobody with the capability to defeat them happened to be looking at you.
You are on one path through a space that contains a very large number of paths. The path you are on is the only one you can see, and it is the only one anyone will ever ask you about.
The question that matters is not what happened. It is what the distribution looks like that this happened to be drawn from.
3.2 The turkey#
Taleb's sharpest illustration of this appears in The Black Swan rather than in Fooled by Randomness, and he traces the underlying point back to Bertrand Russell and, before him, to Hume [2].
A turkey is fed every morning. Each feeding increases its confidence that being fed every morning is what the world does. Its statistical evidence for that belief is at its strongest, its record is at its longest and its data is at its most complete, on the afternoon before Thanksgiving.
The part that should worry anyone running a plant is not the ending. It is the shape of the curve. The turkey's confidence rises monotonically right up to the discontinuity. There is no warning inside the data, because the data is generated by the arrangement that is about to end.
I have never seen a compliance score fall in the quarter before a serious incident. The audit passes. The dashboard is green. The metrics are, if anything, better than they were, because the organization has been working on them. And somewhere in the plant, in the distance between what the reference architecture says is deployed and what is actually installed and connected and running, a path exists that none of the controls were designed to address.
The measurement and the risk are not merely uncorrelated in that period. They move in opposite directions, because effort spent looking good on the measurement is effort not spent on the thing the measurement fails to see.
4. Four illusions I meet in the field#
Right-side thinking, to borrow the phrase Part II will define properly, has recognizable forms. These are the four I encounter most.
4.1 The reference architecture illusion#
There is a diagram on the wall. It shows the network cleanly segmented, the industrial side separated from the corporate side, a demilitarized zone correctly positioned, every connection passing through a controlled chokepoint.
The diagram is accurate. It is an accurate picture of the system as designed.
What is actually running is the product of every operational decision made since commissioning. A vendor access path opened during a maintenance window and never closed. A temporary bridge that became permanent because removing it would require an outage nobody wants to schedule. Equipment that predates the current architecture and was absorbed rather than replaced. Firmware that has not been touched since installation because the process it controls cannot be stopped to update it.
The plant and the drawing diverged years ago, a little at a time, each step defensible on its own.
The audit checked the drawing.
4.2 Compliance read as security posture#
I want to be careful here because this is the point most easily misread.
IEC 62443, NERC CIP, NIS2 and the NIST framework are serious documents. They are built from real incidents and real engineering judgment, and an operator who conforms to them is in a materially better position than one who does not. I have built assessment work on 62443 for years and I would do it again tomorrow.
What they cannot do, by construction, is tell you the risk of a specific facility at a specific moment. A framework is general. It describes what is known to matter across a class of installations. Your exposure is a property of your equipment, your configuration drift, your dependency chain, your people, your adversaries and this week. No general document can carry that, and none of them claim to.
The illusion is not in the framework. It is in the reading. A conformance result is a statement about whether specified controls are present. It is being used as a statement about whether the facility is likely to be compromised. Those are different questions, and the gap between them is where I spend most of my working life.
4.3 Vendor survivorship#
A vendor shows you their reference customers. Every one of them deployed the product and did not have a significant incident.
That list is a sample conditioned on the outcome. It is selected, by definition, from the sites where things went well. The sites that deployed the same product and were compromised anyway are not on it, and they are not on anyone's list, because of the second mechanism in section 2.3. The incident was handled quietly. The customer has no interest in publishing it. The vendor has considerably less.
This is Taleb's fund management market with a different product. You are shown the survivors and invited to read them as proof.
The honest version of that conversation is a question, and it is a fair one to ask: how many deployments of this product have you had, and in how many of them did a significant incident occur anyway. I have asked it. The answers are instructive, including the refusals.
4.4 Fifteen years of doing it this way#
Critical infrastructure runs on institutional memory, and mostly that is a virtue. These are facilities where an untested change can hurt somebody. A bias toward practices that have demonstrably worked is not stubbornness, it is appropriate caution, and I have more sympathy for it than most people in my line of work.
The problem is specific. A practice that has held for fifteen years is evidence about a world that no longer exists. The threat environment is not stationary. The adversaries are better resourced than they were, the software dependency chain is deeper and less visible, the equipment is more connected, and the geopolitical picture that determines who is interested in your sector has changed more than once in that period.
Fifteen unbroken years tells you the practice was adequate against the threats of those years. It is close to silent about the next one.
5. What I am not saying#
Three things, because each of them is a conclusion people reach from this argument and none of them follows.
I am not saying the frameworks should be abandoned. I use them. They encode knowledge that would otherwise have to be rediscovered incident by incident, which is the expensive way.
I am not saying audits are worthless. An audit finds real defects and it finds them cheaply. What it does not do is measure the quantity the operator cares about, and the failure is in treating the one as the other.
I am not saying risk cannot be measured. The whole of this series exists because I think it can. Part III sets out how. What I am saying is that the current measurement is of a different quantity, and that the confidence it generates is not backed by the thing generating it.
The argument is about inference, not about effort. The industry is working hard. It is working hard on the visible side.
6. What the rest of the series does#
Part II, The Two Sides of the Table, states the distinction precisely. One side holds what can be observed: the audit, the dashboard, the unbroken record. The other holds the distribution those observations came from, including the branches that did not run this time. It shows where the drawing and the plant come apart, and states the turkey problem in its proper form as a system whose measured stability rises to the point of transition.
Part III, Running the Paths That Did Not Happen, gives the method. If the risk lives on the unobserved side, the only honest instrument is one that samples it. That part explains what it means to generate a facility's counterfactual histories, what a model has to represent for that sampling to mean anything, and what the method cannot tell you.
Part IV, The Barbell, is the engineering consequence and the most technical of the five. It sets out why enterprise security tooling damages operational technology, mechanism by mechanism, and what the alternative structure looks like.
Part V, Now, Next, Never, is about the decision. If the loss sits in the tail, conventional budgeting allocates against the wrong distribution. That part sets out the allocation that follows and stops short of telling you which supplier to call, which is not my business to say in a paper.
7. References#
[1] N. N. Taleb, Fooled by Randomness: The Hidden Role of Chance in Life and in the Markets. New York: Texere, 2001. Cited for survivorship in populations of identical strategies, the mis-attribution of outcome to skill, and the use of alternative histories as a corrective.
[2] N. N. Taleb, The Black Swan: The Impact of the Highly Improbable. New York: Random House, 2007. Cited for silent evidence and for the turkey, which appears here rather than in [1] and which Taleb traces to Bertrand Russell's chicken and through it to Hume's problem of induction.
[3] International Electrotechnical Commission, IEC 62443, Industrial communication networks, Security for industrial automation and control systems. Cited as a framework in general use, including by the author, and not as a target of this argument.
[4] National Institute of Standards and Technology, Framework for Improving Critical Infrastructure Cybersecurity. Cited in the same sense as [3].
[5] Directive (EU) 2022/2555 on measures for a high common level of cybersecurity across the Union, NIS2. Cited in the same sense as [3].
[6] North American Electric Reliability Corporation, Critical Infrastructure Protection reliability standards. Cited in the same sense as [3].
Note on this bibliography#
The field observations in sections 3 and 4 are the author's own, drawn from assessment and engineering work in the energy, rail, water and manufacturing sectors in North America, Australasia and Europe. They are offered as testimony and are not supported by a citation, because no published source records them. Where this series makes a claim that a published source does support, the source is named. Where it makes a claim resting on the author's own observation, it says so in the sentence rather than in a footnote.